A method and system for adjusting shooting lens based on machine vision
Through the machine vision-based shooting lens adjustment method, the video data of multiple lenses is used to detect and feature extraction, and the optimal lens is selected for picture synthesis, which solves the problems of frame dropping and inefficiency in real-time shooting scenes, and achieves high-quality and complete panoramic video generation.
Patent Information
- Application Number
- CN202510017045.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art is prone to frame loss in real-time shooting scenes, and the shooting efficiency and quality are low.
Using a shooting lens adjustment method based on machine vision, by obtaining video data of multiple lenses, detecting targets based on timestamps and extracting feature information, generating an objective function to select the optimal lens, and performing picture synthesis to generate panoramic video.
It significantly improves the video quality and integrity of video processing from multi-lens perspectives, provides a clearer and smoother visual experience, and avoids frame loss problems.
Smart Images

Figure CN119450228B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of shooting lenses, and in particular to a shooting lens adjustment method and system based on machine vision. Background Art
[0002] The technical background of the camera lens adjustment method based on machine vision mainly stems from the demand for high-precision and high-stability imaging of the machine vision system. Machine vision, as a technology that uses machines to replace human eyes for image recognition, measurement and judgment, is widely used in industrial automation, quality inspection, intelligent monitoring and other fields. In order to ensure that the machine vision system can complete the task accurately and efficiently, the adjustment of the camera lens is particularly critical.
[0003] In the prior art, the image captured by the camera lens is generally detected to control the shooting mode of the lens accordingly. However, this method is often not suitable for real-time shooting scenes, and is prone to frame loss and has relatively low shooting efficiency and quality. Summary of the invention
[0004] The embodiments of the present application provide a method and system for adjusting a shooting lens based on machine vision, which can at least to some extent solve the problem of easy frame loss and relatively low shooting efficiency and shooting quality.
[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0006] According to one aspect of the present application, a method for adjusting a shooting lens based on machine vision is provided, which is applied to a shooting device, wherein the shooting device includes a plurality of lenses, and the method includes:
[0007] Acquire video data captured by the lens;
[0008] Based on the timestamps of the video data, detecting the target in each of the video data, and extracting the feature information corresponding to the target at each timestamp;
[0009] Generate an objective function according to the feature information corresponding to each of the video data, and select a shot corresponding to the optimal objective function as the target shot at the current moment;
[0010] The video data collected by the target lens is used as the main data, and the video data collected by the lens other than the target lens in the shooting device is used as the backup data, and the picture synthesis is performed to generate a panoramic picture, and then the target lens corresponding to the next timestamp is iteratively selected in the same way and the picture synthesis is performed until a panoramic video based on the timestamp is generated;
[0011] The panoramic video is transmitted to a display device for display.
[0012] In the present application, based on the aforementioned scheme, the targets in each of the video data are detected based on the timestamp of the video data, and the feature information corresponding to the targets at each timestamp is extracted, including: based on the timestamp of the video data, the video data is reconstructed to generate a reconstructed image; the targets in the reconstructed image are detected, and the feature information corresponding to the targets at each timestamp is extracted; based on a preset data structure, each target and its feature information at each timestamp is stored.
[0013] In the present application, based on the aforementioned scheme, the video data is reconstructed based on the timestamp of the video data to generate a reconstructed image, including: decoding the video data from a compressed format into an original frame sequence through video decoding, and extracting the timestamp of each frame; extracting a video image from the original frame sequence corresponding to each timestamp; and reconstructing the video image pixel by pixel for a pixel block in the video image to generate a reconstructed image.
[0014] In the present application, based on the aforementioned scheme, the objective function is generated according to the feature information corresponding to each of the video data, and the lens corresponding to the optimal objective function is selected as the target lens at the current moment, including: generating the objective function based on the feature information extracted from the i-th lens and the j-th lens at timestamp t, and selecting the lens corresponding to the optimal objective function as the target lens at the current moment.
[0015] In the present application, based on the aforementioned scheme, the video data collected by the target lens is used as the main data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data, and the picture synthesis is performed to generate a panoramic picture, including: the video data collected by the target lens is used as the main data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data; based on the main data, the backup data is filled into the main data to generate a panoramic picture.
[0016] In the present application, based on the above-mentioned solution, it also includes: if it is detected that a picture in the main data is missing, an area corresponding to the missing part is selected from the backup data corresponding to the main data to fill it.
[0017] In the present application, based on the aforementioned solution, the transmitting the panoramic video to a display device for display includes: connecting a shooting device and a display device via a transmission device; transmitting the panoramic video to the display device for display.
[0018] According to one aspect of the present application, a shooting lens adjustment system based on machine vision is provided, which is applied to a shooting device, wherein the shooting device includes a plurality of lenses, including:
[0019] An acquisition unit, used to acquire video data captured by the lens;
[0020] An extraction unit, configured to detect a target in each of the video data based on the timestamps of the video data, and extract feature information corresponding to the target at each timestamp;
[0021] A target unit, used to generate a target function according to the feature information corresponding to each of the video data, and select a shot corresponding to the optimal target function as the target shot at the current moment;
[0022] A synthesis unit, used to use the video data collected by the target lens as main data and the video data collected by lenses other than the target lens in the shooting device as backup data, perform picture synthesis to generate a panoramic picture, and then iteratively select the target lens corresponding to the next timestamp in the same way and perform picture synthesis until a panoramic video based on the timestamp is generated;
[0023] The display unit is used to transmit the panoramic video to a display device for display.
[0024] In the present application, based on the aforementioned scheme, the targets in each of the video data are detected based on the timestamp of the video data, and the feature information corresponding to the targets at each timestamp is extracted, including: based on the timestamp of the video data, the video data is reconstructed to generate a reconstructed image; the targets in the reconstructed image are detected, and the feature information corresponding to the targets at each timestamp is extracted; based on a preset data structure, each target and its feature information at each timestamp is stored.
[0025] In the present application, based on the aforementioned scheme, the video data is reconstructed based on the timestamp of the video data to generate a reconstructed image, including: decoding the video data from a compressed format into an original frame sequence through video decoding, and extracting the timestamp of each frame; extracting a video image from the original frame sequence corresponding to each timestamp; and reconstructing the video image pixel by pixel for a pixel block in the video image to generate a reconstructed image.
[0026] In the present application, based on the aforementioned scheme, the objective function is generated according to the feature information corresponding to each of the video data, and the lens corresponding to the optimal objective function is selected as the target lens at the current moment, including: generating the objective function based on the feature information extracted from the i-th lens and the j-th lens at timestamp t, and selecting the lens corresponding to the optimal objective function as the target lens at the current moment.
[0027] In the present application, based on the aforementioned scheme, the video data collected by the target lens is used as the main data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data, and the picture synthesis is performed to generate a panoramic picture, including: the video data collected by the target lens is used as the main data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data; based on the main data, the backup data is filled into the main data to generate a panoramic picture.
[0028] In the present application, based on the above-mentioned solution, it also includes: if it is detected that a picture in the main data is missing, an area corresponding to the missing part is selected from the backup data corresponding to the main data to fill it.
[0029] In the present application, based on the aforementioned solution, the transmitting the panoramic video to a display device for display includes: connecting a shooting device and a display device via a transmission device; transmitting the panoramic video to the display device for display.
[0030] According to one aspect of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a shooting lens adjustment method based on machine vision as described in the above embodiment is implemented.
[0031] According to one aspect of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a shooting lens adjustment method based on machine vision as described in the above-mentioned embodiments.
[0032] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes a method for adjusting a shooting lens based on machine vision provided in the above-mentioned various optional implementations.
[0033] In the technical solution of this application, on the one hand, by integrating multiple shooting lenses, machine vision technology is used to capture and process video data in real time, ensuring the timeliness of video data acquisition. Target detection and feature extraction based on timestamp synchronization ensure the precise alignment of each lens data and the efficient capture of key information.
[0034] On the other hand, by intelligently evaluating the feature information captured by each lens, the main data and backup data are determined, and the backup lens data is used to fill in the missing or optimize the picture. The video data of the next timestamp is then iteratively processed to achieve virtual adjustment and seamless connection between the optimal lens and the backup lens, which significantly improves the video quality and integrity of video processing under multi-lens perspectives and provides a clearer and smoother visual experience.
[0035] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0037] Figure 1 The flowchart of a method for adjusting a shooting lens based on machine vision in one embodiment of the present application is schematically shown.
[0038] Figure 2 The flowchart of extracting feature information in one embodiment of the present application is schematically shown.
[0039] Figure 3 A schematic diagram of a shooting lens adjustment system based on machine vision in an embodiment of the present application is schematically shown.
[0040] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0041] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.
[0042] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0043] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0044] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0045] The implementation details of the technical solution of this application are described in detail below:
[0046] Figure 1 FIG. 1 is a flowchart of a method for adjusting a shooting lens based on machine vision according to an embodiment of the present application. Figure 1 As shown, the shooting lens adjustment method based on machine vision at least includes steps S110 to S150, which are described in detail as follows:
[0047] In step S110, video data captured by the lens is obtained.
[0048] In one embodiment of the present application, the present invention is applied to a shooting device, wherein the shooting device includes multiple lenses. The multiple lenses are used to simultaneously acquire video data of a target object, wherein the target object may be a precision instrument, an operating platform, and the like.
[0049] In one embodiment of the present application, it can be applied to the production process monitoring of the equipment, and the production process of the equipment can be collected through multiple lenses to perform multi-lens and blind-angle production monitoring. It can also be applied to the video shooting of remote surgery and surgical teaching to achieve the same purpose.
[0050] Video data is captured simultaneously by multiple lenses to ensure the real-time and accuracy of video data, providing a rich data source for subsequent target detection, feature extraction and lens selection.
[0051] In step S120, based on the timestamps of the video data, the target in each of the video data is detected, and the feature information corresponding to the target at each timestamp is extracted.
[0052] like Figure 2 As shown, in one embodiment of the present application, based on the timestamps of the video data, detecting the target in each of the video data, and extracting the feature information corresponding to the target at each timestamp, includes:
[0053] S210, reconstructing the video data based on the timestamp of the video data to generate a reconstructed image;
[0054] S220, detecting a target in the reconstructed image, and extracting feature information corresponding to the target at each time stamp;
[0055] S230, based on a preset data structure, storing each target and its characteristic information at each timestamp.
[0056] In one embodiment of the present application, in step S210, reconstructing the video data based on the timestamp of the video data to generate a reconstructed image includes:
[0057] Decode the video data from the compressed format into the original frame sequence through video decoding, and extract the timestamp of each frame;
[0058] Extracting video images from the original frame sequences corresponding to the timestamps;
[0059] For the pixel blocks in the video image, the video image is pixel reconstructed to generate a reconstructed image.
[0060] In one embodiment of the present application, the video is usually stored in a compressed format to save storage space and improve transmission efficiency. Video decoding is the process of converting these compressed video data into an original, uncompressed frame sequence. During the decoding process, the encoded data of the video file is read by a video decoder and decoded into a series of continuous image frames.
[0061] At the same time, the decoder also extracts the timestamp information of each frame for subsequent frame processing and synchronization. In the decoded original frame sequence, each frame has a timestamp associated with it. Based on the timestamp, the video image of a specific time point or time period can be extracted for video editing, analysis or extraction of specific frames.
[0062] In the extracted video image, the image is divided into multiple pixel blocks, including local areas for compression, encoding or processing of the image. These pixel blocks are processed such as denoising, enhancement, compression, etc. Pixel reconstruction is to perform some form of transformation or optimization on the pixel block or the entire image to generate a new image with different characteristics (such as resolution, color, etc.).
[0063] Specifically, the video image is divided into a plurality of small pixel blocks according to a preset number of pixels, each block containing a certain number of pixels. Each pixel block is subjected to pixel transformation to generate the transformation coefficient corresponding to each pixel point. for:
[0064]
[0065] in, is the pixel value of the original pixel block, Represents the horizontal and vertical coordinates of the center of the pixel block, N is the pixel block size, and Corresponding to the scaling factors on the horizontal and vertical axes respectively, cos is the cosine function, sin is the sine function, It is pi.
[0066] After the transform coefficients are generated based on the pixel blocks, the video image is pixel-reconstructed based on the transform coefficients. This may be a method of linearly processing the pixel points in the video image according to the transform coefficients to generate a reconstructed image.
[0067] Furthermore, after the reconstructed image is generated, the quality of the reconstructed image is evaluated. and original video image , generating the difference parameters for:
[0068]
[0069] in, M, N is the number of pixels in the horizontal and vertical directions of the video image. The original video image is in ( ) pixel value at the position, is the reconstructed image in ( ) is the pixel value at the position.
[0070] After the difference parameter is calculated, it is compared with the set threshold. If it is greater than or equal to the set threshold, it means that the difference with the original image is high and may have been distorted, so image reconstruction is required. Through the video decoding and reconstruction process, the compressed video data can be restored to the original frame sequence to ensure that the image quality is not affected by compression.
[0071] After generating the reconstructed image, target detection is performed, for example, by extracting image features through a convolutional neural network and using these features to identify the target. Appropriate feature types are selected according to task requirements, such as color features, texture features, shape features, and motion features.
[0072] For each detected target, feature information is extracted using a corresponding feature extraction algorithm. For example, color features can be extracted using a color histogram, shape features can be extracted using contour analysis or shape context, and motion features can be extracted using trajectory tracking. This embodiment provides an accurate data basis for subsequent shot selection and picture synthesis by extracting the timestamp of each frame and target feature information in the reconstructed image.
[0073] Store the extracted feature information in an appropriate data structure, such as a database, file system, or memory cache. Create an index for the stored feature information to quickly query and access the target features at a specific timestamp. The preset data structure can efficiently store and manage the target and its feature information at each timestamp, which is convenient for subsequent processing.
[0074] The above process uses timestamp information to ensure the consistency of video data captured by different lenses in the time dimension, which is convenient for subsequent picture synthesis. Object detection and feature extraction technology can accurately identify key objects in the video and extract their important features, providing a basis for subsequent lens selection.
[0075] In step S130, an objective function is generated according to the feature information corresponding to each of the video data, and a shot corresponding to the optimal objective function is selected as the target shot at the current moment.
[0076] In one embodiment of the present application, based on i The lens and j Shots at timestamp t The feature information extracted at the time is used to generate the objective function, and the lens corresponding to the optimal objective function is selected as the target lens at the current moment.
[0077] In one embodiment of the present application, based on i The lens and j Shots at timestamp t Feature information extracted at all times to construct the objective function for:
[0078]
[0079] in, represents the time factor generated based on historical data, Respectively represent i The lens andj Shots at timestamp t The feature information extracted when Indicates i The lens and j The correlation factor between the lenses, k is the total number of lenses, and log is a logarithmic function.
[0080] Based on the feature information between the lenses, the differences between the video data captured by different lenses are quantified. By solving the minimization problem of the objective function, the lens corresponding to the optimal objective function is selected as t The target shot corresponding to the moment.
[0081] In the calculation t Moments after the target shot, reacquire t +1 The video data collected at time 1 is used for a new round of calculation and determination t The target lens and spare lens corresponding to the +1 moment undergo relevant image processing.
[0082] The above process can improve the quality of image acquisition and processing effect. Constructing the objective function and selecting the lens corresponding to the optimal objective function as the target lens at the current moment can ensure that the selected lens performs best in capturing the target and improve the video quality.
[0083] In step S140, the video data captured by the target lens is used as the main data, and the video data captured by the lenses other than the target lens in the shooting device is used as the backup data, and the pictures are synthesized to generate a panoramic picture. Then, the target lens corresponding to the next timestamp is iteratively selected and the pictures are synthesized in the same way until a panoramic video based on the timestamp is generated.
[0084] In one embodiment of the present application, the video data collected by the target lens is used as the primary data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data. The primary data usually has higher quality, more complete viewing angle or more accurate target information. The video data collected by the target lens is used as the primary data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data, which can fully utilize the advantages of multiple lenses to generate a more complete and accurate panoramic video.
[0085] In one embodiment of the present application, the video data collected by the target lens is used as the main data, and the video data collected by the lens other than the target lens in the shooting device is used as the backup data, and the picture synthesis is performed to generate a panoramic picture, including:
[0086] The video data collected by the target lens is used as the main data, and the video data collected by the lens other than the target lens in the shooting device is used as the backup data;
[0087] Based on the main data, the backup data is filled into the main data to generate a panoramic picture.
[0088] Optionally, deep learning algorithms can be used for image fusion during the synthesis process. By training a neural network model, it learns how to fill the backup data into the main data based on the main data to generate a high-quality fused image. This method has good robustness when dealing with complex scenes and changing lighting conditions, and can ensure the continuity of panoramic videos in time and space.
[0089] Specifically, in this embodiment, after the main data and spare data in the video data corresponding to the current timestamp are synthesized to generate a panoramic picture, each subsequent timestamp needs to be processed in the same way, that is, target detection, feature information extraction, and target lens selection are performed on the video data corresponding to each timestamp in the video sequence (that is, "the video data collected by the target lens is used as the main data, and the video data collected by the lenses other than the target lens in the shooting device is used as the spare data, and the pictures are synthesized to generate a panoramic picture" in step S120, step S130, and step S140 are executed once for each timestamp). In this way, adaptive adjustment between the target lens and the remaining lenses based on the timestamp is achieved, and then the synthesis of the main data and spare data based on the timestamp is achieved to obtain the panoramic picture corresponding to each timestamp. Afterwards, a panoramic video based on the timestamp is generated based on the playback order, and the panoramic picture can also be played based on the timestamp while the video data is processed. In this embodiment, the target lens is adaptively adjusted and the main data and spare data are synthesized at the same time, so as to achieve virtual adjustment and seamless connection of panoramic videos under multiple lenses.
[0090] In one embodiment of the present application, it further includes: if it is detected that a picture in the main data is missing, an area corresponding to the missing part is selected from the spare data corresponding to the main data for filling.
[0091] In one embodiment of the present application, for each timestamp, the best picture is selected from the primary data and the backup data for synthesis. This can be achieved by comparing the characteristic information (such as position, size, clarity, etc.) of the target in different pictures. If there is a problem with the picture in the primary data at a certain timestamp (such as occlusion, blur, etc.), an alternative picture can be selected from the backup data.
[0092] When a missing image is detected in the main data, the corresponding area can be selected from the backup data to fill in, ensuring the integrity and continuity of the panoramic video. This improves the robustness and reliability of the video data and reduces the degradation of video quality caused by missing data.
[0093] Optionally, if there are problems with the pictures in both the primary data and the backup data at a certain timestamp, an interpolation algorithm can be used to generate a substitute picture, or similar pictures can be borrowed from other timestamps to fill the gap.
[0094] Optionally, the synthesized panoramic images can be spliced in time order to generate a panoramic video. In order to maintain the fluency and continuity of the video, operations such as inter-frame interpolation and smoothing can also be performed.
[0095] In summary, this method can generate high-quality panoramic video through the collaborative work of multiple lenses, target detection and feature extraction, lens selection, picture synthesis and other steps, and transmit it to the display device for display. This method has broad application prospects in video surveillance, live sports events, virtual reality, live surgery and other fields.
[0096] In step S150, the panoramic video is transmitted to a display device for display.
[0097] In one embodiment of the present application, the shooting device is connected to the display device via a transmission device; the panoramic video is transmitted to the display device for display. By connecting the shooting device to the display device via the transmission device, the generated panoramic video can be transmitted to the display device in real time for display. This provides an intuitive and clear video display effect, enhancing the user's viewing experience.
[0098] In one embodiment of the present application, the camera and the display device are connected via a transmission device. Ensure that the connection is firm to avoid signal loss or interruption. If wireless network transmission or video streaming transmission is adopted, corresponding software configuration is performed on the camera and the display device, such as setting the IP address, port number, encoding format, etc.
[0099] Optionally, during the display process, the panoramic video is previewed on the display device to check whether the image quality, audio synchronization, frame rate, etc. are normal.
[0100] Optionally, adjust video parameters or display settings to achieve the best display effect, ensuring that the panoramic video is displayed in full screen on the display device to fully utilize the display area and the audience's perspective.
[0101] Optionally, for panoramic videos supporting 360-degree viewing angles, special playback software or hardware may be configured to achieve an all-round viewing experience.
[0102] In the technical solution of the present application, the video data captured by the lens is obtained; based on the timestamp of the video data, the target in each of the video data is detected, and the feature information corresponding to the target at each timestamp is extracted; according to the feature information corresponding to each of the video data, an objective function is generated, and the lens corresponding to the optimal objective function is selected as the target lens at the current moment; the video data captured by the target lens is used as the main data, and the video data captured by the lens other than the target lens in the shooting device is used as the backup data, and the picture is synthesized to generate a panoramic picture, and then the target lens corresponding to the next timestamp is iteratively selected in the same way and the picture is synthesized until a panoramic video based on the timestamp is generated; the panoramic video is transmitted to a display device for display. By integrating multiple shooting lenses, machine vision technology is used to capture and process video data in real time. Target detection and feature extraction based on timestamp synchronization ensure the precise alignment of each lens data and the efficient capture of key information. By intelligently evaluating the feature information captured by each lens, determining the main data and backup data, and using the backup lens data to fill in missing or optimize the picture, and then iteratively processing the video data of the next timestamp, virtual adjustment and seamless connection between the optimal lens and the backup lens are achieved, significantly improving the video quality and integrity of video processing under multi-lens perspectives, and providing a clearer and smoother visual experience.
[0103] The following describes an embodiment of the device of the present application, which can be used to execute a method for adjusting a shooting lens based on machine vision in the above-mentioned embodiment of the present application. It can be understood that the device can be a computer program (including program code) running in a computer device, for example, the device is an application software; the device can be used to execute the corresponding steps in the method provided in the embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method for adjusting a shooting lens based on machine vision in the above-mentioned embodiment of the present application.
[0104] Figure 3 A block diagram of a shooting lens adjustment system based on machine vision according to an embodiment of the present application is shown.
[0105] Reference Figure 3 As shown, according to an embodiment of the present application, a shooting lens adjustment system based on machine vision includes:
[0106] An acquisition unit 310 is used to acquire video data captured by the lens;
[0107] An extraction unit 320, configured to detect a target in each of the video data based on the timestamps of the video data, and extract feature information corresponding to the target at each timestamp;
[0108] A target unit 330 is used to generate a target function according to the feature information corresponding to each of the video data, and select a shot corresponding to the optimal target function as a target shot at a current moment;
[0109] A synthesis unit 340 is used to use the video data collected by the target lens as primary data and the video data collected by lenses other than the target lens in the shooting device as backup data to perform picture synthesis to generate a panoramic picture, and then iteratively select the target lens corresponding to the next timestamp in the same way and perform picture synthesis until a panoramic video based on the timestamp is generated;
[0110] The display unit 350 is used to transmit the panoramic video to a display device for display.
[0111] This embodiment intelligently evaluates the feature information captured by each lens to determine the main data and backup data, so as to use the backup lens data to fill in the missing or optimize the picture, and then iteratively processes the video data of the next timestamp to achieve virtual adjustment and seamless connection between the optimal lens and the backup lens, thereby significantly improving the video quality and integrity of video processing under multi-lens perspectives and providing a clearer and smoother visual experience.
[0112] In the present application, based on the aforementioned scheme, the targets in each of the video data are detected based on the timestamp of the video data, and the feature information corresponding to the targets at each timestamp is extracted, including: based on the timestamp of the video data, the video data is reconstructed to generate a reconstructed image; the targets in the reconstructed image are detected, and the feature information corresponding to the targets at each timestamp is extracted; based on a preset data structure, each target and its feature information at each timestamp is stored.
[0113] In the present application, based on the aforementioned scheme, the video data is reconstructed based on the timestamp of the video data to generate a reconstructed image, including: decoding the video data from a compressed format into an original frame sequence through video decoding, and extracting the timestamp of each frame; extracting a video image from the original frame sequence corresponding to each timestamp; and reconstructing the video image pixel by pixel for a pixel block in the video image to generate a reconstructed image.
[0114] In the present application, based on the above-mentioned solution, generating an objective function according to the feature information corresponding to each of the video data, and selecting the lens corresponding to the optimal objective function as the target lens at the current moment, includes: based on the first i The lens and j Shots at timestamp t The feature information extracted at the time is used to generate the objective function, and the lens corresponding to the optimal objective function is selected as the target lens at the current moment.
[0115] In the present application, based on the aforementioned scheme, the video data collected by the target lens is used as the main data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data, and the picture synthesis is performed to generate a panoramic picture, including: the video data collected by the target lens is used as the main data, and the video data collected by the lenses other than the target lens in the shooting device is used as the backup data; based on the main data, the backup data is filled into the main data to generate a panoramic picture.
[0116] In the present application, based on the above-mentioned solution, it also includes: if it is detected that a picture in the main data is missing, an area corresponding to the missing part is selected from the backup data corresponding to the main data to fill it.
[0117] In the present application, based on the aforementioned solution, the transmitting the panoramic video to a display device for display includes: connecting a shooting device and a display device via a transmission device; transmitting the panoramic video to the display device for display.
[0118] In the technical solution of the present application, the video data captured by the lens is obtained; based on the timestamp of the video data, the target in each of the video data is detected, and the feature information corresponding to the target at each timestamp is extracted; according to the feature information corresponding to each of the video data, an objective function is generated, and the lens corresponding to the optimal objective function is selected as the target lens at the current moment; the video data captured by the target lens is used as the main data, and the video data captured by the lens other than the target lens in the shooting device is used as the spare data, and the picture is synthesized to generate a panoramic picture, and then the target lens corresponding to the next timestamp is iteratively selected in the same way and the picture is synthesized until a panoramic video based on the timestamp is generated; the panoramic video is transmitted to a display device for display. By integrating multiple shooting lenses, machine vision technology is used to capture and process video data in real time. Target detection and feature extraction based on timestamp synchronization ensure the precise alignment of each lens data and the efficient capture of key information. By intelligently evaluating the feature information captured by each lens, determining the main data and spare data, and using the spare lens data to fill in the missing or optimize the picture, the accurate synthesis of the panoramic video is achieved. It not only significantly improves the video quality and integrity, but also enhances the robustness and real-time performance of the video, providing a clearer and smoother visual experience.
[0119] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.
[0120] It should be noted that the computer system of the electronic device in this embodiment is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0121] In this embodiment, the computer system includes a central processing unit 401, which can perform various appropriate actions and processes according to the program stored in the read-only memory 402 or the program loaded from the storage part 408 to the random access memory 403, such as executing the method described in the above embodiment. Various programs and data required for system operation are also stored in the random access memory 403. The central processing unit 401, the read-only memory 402 and the random access memory 403 are connected to each other through a bus 404. The input / output interface 405 is also connected to the bus 404.
[0122] The following components are connected to the input / output interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read therefrom is installed into the storage section 408 as needed.
[0123] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 409, and / or installed from a removable medium 411. When the computer program is executed by the central processing unit 401, various functions defined in the system of the present application are executed.
[0124] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0125] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0126] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.
[0127] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above-mentioned various optional implementations.
[0128] This embodiment intelligently evaluates the feature information captured by each lens to determine the main data and backup data, so as to use the backup lens data to fill in the missing or optimize the picture, and then iteratively processes the video data of the next timestamp to achieve virtual adjustment and seamless connection between the optimal lens and the backup lens, thereby significantly improving the video quality and integrity of video processing under multi-lens perspectives and providing a clearer and smoother visual experience.
[0129] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.
[0130] This embodiment intelligently evaluates the feature information captured by each lens to determine the main data and backup data, so as to use the backup lens data to fill in the missing or optimize the picture, and then iteratively processes the video data of the next timestamp to achieve virtual adjustment and seamless connection between the optimal lens and the backup lens, thereby significantly improving the video quality and integrity of video processing under multi-lens perspectives and providing a clearer and smoother visual experience.
[0131] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0132] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software, or by combining software with necessary hardware. Therefore, the technical solution according to the implementation method of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation method of the present application.
[0133] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.
[0134] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for adjusting a shooting lens based on machine vision, characterized in that: Applied to a photographing device, the photographing device includes a plurality of lenses, and the method includes: Acquire video data captured by the lens; Based on the timestamps of the video data, detecting the target in each of the video data, and extracting the feature information corresponding to the target at each timestamp; Generate an objective function according to the feature information corresponding to each of the video data, and select a shot corresponding to the optimal objective function as the target shot at the current moment; The video data collected by the target lens is used as the main data, and the video data collected by the lens other than the target lens in the shooting device is used as the backup data, and the picture is synthesized to generate a panoramic picture, and then the target lens corresponding to the next timestamp is iteratively selected in the same way and the picture is synthesized until a panoramic video based on the timestamp is generated; Transmitting the panoramic video to a display device for display; The step of generating an objective function according to the feature information corresponding to each of the video data and selecting a shot corresponding to the optimal objective function as the target shot at the current moment includes: Based on i The lens and j Shots at timestamp t The feature information extracted during the experiment is used to construct the objective function for: in, represents the time factor generated based on historical data, Respectively represent i The lens and j Shots at timestamp t The feature information extracted when Indicates i The lens and j The correlation factor between the lenses, k is the total number of lenses, log is the logarithmic function; Select the lens corresponding to the optimal objective function as t The target shot corresponding to the moment.
2. The method for adjusting a shooting lens based on machine vision according to claim 1, characterized in that: Detecting a target in each of the video data based on the timestamps of the video data, and extracting feature information corresponding to the target at each timestamp, including: Reconstructing the video data based on the timestamp of the video data to generate a reconstructed image; Detecting a target in the reconstructed image and extracting feature information corresponding to the target at each time stamp; Based on the preset data structure, each target and its feature information at each timestamp is stored.
3. The method for adjusting a shooting lens based on machine vision according to claim 2, characterized in that: Reconstructing the video data based on the timestamp of the video data to generate a reconstructed image includes: Decode the video data from the compressed format into the original frame sequence through video decoding, and extract the timestamp of each frame; Extracting video images from the original frame sequences corresponding to the timestamps; Based on the pixel blocks in the video image, the video image is pixel reconstructed to generate a reconstructed image.
4. The method for adjusting a shooting lens based on machine vision according to claim 1, characterized in that: The method uses the video data collected by the target lens as main data and the video data collected by lenses other than the target lens in the shooting device as backup data to perform picture synthesis to generate a panoramic picture, including: The video data collected by the target lens is used as the main data, and the video data collected by the lens other than the target lens in the shooting device is used as the backup data; Based on the main data, the backup data is filled into the main data to generate a panoramic picture.
5. The method for adjusting a shooting lens based on machine vision according to claim 4, characterized in that: Also includes: If it is detected that a picture in the main data is missing, an area corresponding to the missing part is selected from the spare data corresponding to the main data to fill it.
6. The method for adjusting a shooting lens based on machine vision according to claim 1, characterized in that: Transmitting the panoramic video to a display device for display, including: Connecting the shooting device to the display device via a transmission device; The panoramic video is transmitted to a display device for display.
7. A shooting lens adjustment system based on machine vision, characterized in that: Applied to a shooting device, the shooting device includes multiple lenses, including: An acquisition unit, used to acquire video data captured by the lens; An extraction unit, configured to detect a target in each of the video data based on the timestamps of the video data, and extract feature information corresponding to the target at each timestamp; A target unit, used to generate a target function according to the feature information corresponding to each of the video data, and select a shot corresponding to the optimal target function as the target shot at the current moment; A synthesis unit, used to use the video data collected by the target lens as main data and the video data collected by lenses other than the target lens in the shooting device as backup data, perform picture synthesis to generate a panoramic picture, and then iteratively select the target lens corresponding to the next timestamp in the same way and perform picture synthesis until a panoramic video based on the timestamp is generated; A display unit, used for transmitting the panoramic video to a display device for display; The step of generating an objective function according to the feature information corresponding to each of the video data and selecting a shot corresponding to the optimal objective function as the target shot at the current moment includes: Based on i The lens and j Shots at timestamp t The feature information extracted during the experiment is used to construct the objective function for: in, represents the time factor generated based on historical data, Respectively represent i The lens and j Shots at timestamp t The feature information extracted when Indicates i The lens and j The correlation factor between the lenses, k is the total number of lenses, log is the logarithmic function; Select the lens corresponding to the optimal objective function as t The target shot corresponding to the moment.
8. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a shooting lens adjustment method based on machine vision as claimed in any one of claims 1 to 7 is implemented.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement a machine vision-based shooting lens adjustment method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Video follow-up-based multi-operation-area safety control method and video follow-up-based multi-operation-area safety control system
CN118018681A