Machine vision-based probe image stitching method, device, medium and product
Patent Information
- Application Number
- CN202611275502.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-21
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]上述基于显式几何标定的图像拼接方法在处理超广角镜头(水平视场角≥180°)时存在固有缺陷:由于超广角镜头边缘存在严重的非线性径向畸变,传统方法依赖相机内参、外参及畸变模型进行显式投影变换,该过程会导致边缘区域像素信息截断、重采样失真或有效视野过度裁剪,造成拼接后图像边缘信息严重缺失,图像拼接效果较差
[0024]第五方面,本申请实施例提供一种计算机程序产品,所述计算机程序产品包括计算机程序,所述计算机程序被处理器执行时实现第一方面或第一方面的任意一种可能的实现方式提供的方法。
Smart Images

Figure CN122820437A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image data processing technology, and more specifically, to a machine vision-based detection image stitching method, device, medium, and product. Background Technology
[0002] In perimeter security and other surveillance scenarios, to eliminate blind spots caused by limitations in camera installation height and angle, a multi-camera node deployment method is often used (such as deploying ultra-wide-angle cameras equipped with vibration detectors at intervals of 5-6 meters), and image stitching technology is used to synthesize multiple video streams into a panoramic image. Traditional image stitching methods typically follow a multi-stage sequential process of feature extraction, feature matching, geometric transformation estimation, image registration, and fusion optimization.
[0003] The image stitching method based on explicit geometric calibration has inherent defects when dealing with ultra-wide-angle lenses (horizontal field of view ≥180°): due to the severe nonlinear radial distortion at the edges of ultra-wide-angle lenses, traditional methods rely on camera intrinsic parameters, extrinsic parameters and distortion models to perform explicit projection transformation. This process can lead to truncation of pixel information in edge regions, resampling distortion or excessive cropping of the effective field of view, resulting in severe loss of edge information in the stitched image and poor image stitching effect. Summary of the Invention
[0004] The purpose of this application is to provide a machine vision-based detection image stitching method, device, medium, and product to solve the above-mentioned problems.
[0005] In a first aspect, embodiments of this application provide a machine vision-based detection image stitching method, comprising: acquiring multiple detection images to be stitched from multiple perspectives; performing feature encoding on each of the detection images to be stitched, and embedding perspective identification information during the encoding process to obtain feature representations of each perspective; performing cross-perspective feature fusion based on the correlation between the feature representations of different perspectives to obtain fused global features; and decoding and generating a panoramic image from the global features based on a learnable query vector.
[0006] In the implementation of the above scheme, an end-to-end deep learning architecture replaces the traditional multi-stage geometric calibration and projection transformation process. Throughout the feature encoding, cross-viewpoint fusion, and panoramic generation processes, there is no need to rely on explicit calculations of camera intrinsic and extrinsic parameters and distortion models. Panoramic results are directly generated from the original distorted image, avoiding edge information truncation and resampling distortion caused by explicit projection transformation, thus achieving complete preservation of edge region information in ultra-wide-angle lenses. Furthermore, by embedding viewpoint identification information during feature encoding, the physical layout of the camera array is encoded a priori into a form that the model can perceive. This allows the model to understand the spatial topological relationships and viewpoint identities of each input image even without calibration. This improves the accuracy and convergence stability of multi-view feature alignment. Furthermore, by establishing the correlation between feature representations from different perspectives and performing cross-view feature fusion, global context modeling is used to replace the traditional local feature point matching mechanism, effectively handling parallax misalignment, dynamic occlusion, and exposure differences, thus improving the robustness and visual consistency of the stitching results in complex scenes. On another front, a learnable query vector is used to directly decode and generate panoramic images from the fused global features. Adaptive feature sampling replaces the traditional back projection mapping and interpolation fusion process, avoiding interpolation holes, black border cropping, and ghosting issues, and achieving direct mapping from feature space to image space.
[0007] In one implementation of the first aspect, the step of decoding and generating a panoramic image from the global features based on learnable query vectors includes: dynamically predicting the sampling position and offset of each learnable query vector in the global feature space; sampling and aggregating feature information from the global features based on the sampling position and the offset to generate corresponding panoramic image pixels; wherein each query vector is configured to generate the panoramic image pixels at the corresponding spatial position in the panoramic image.
[0008] In the implementation of the above scheme, panoramic image pixels are generated directly from global features through learnable query vectors, avoiding the back projection mapping and interpolation filling process based on geometric transformation in traditional methods. This eliminates interpolation holes, black edge cropping, and resampling distortion caused by discretization sampling, and realizes direct mapping from feature space to image space, improving the integrity of panoramic image generation. On the other hand, by dynamically predicting the sampling position and offset of the query vector in feature space, each query only aggregates the local feature information most relevant to the current output position, avoiding the detail blurring caused by global average pooling. This reduces computational overhead while enhancing the ability to preserve edges and high-frequency details. Furthermore, by configuring each query vector to generate pixels at specific spatial locations in the panoramic image, a fixed mapping relationship between query identity and output position is established, giving the decoding process a clear spatial responsibility division. This avoids position confusion after multi-view feature fusion and improves the spatial consistency and structural accuracy of panoramic image generation.
[0009] In one implementation of the first aspect, the cross-perspective feature fusion based on the correlation between the feature representations from different perspectives to obtain the fused global features includes: determining a fusion mode for cross-perspective feature fusion based on system performance requirements; wherein the system performance requirements include high precision requirements and real-time requirements; the fusion mode includes a global fusion mode matching the high precision requirements and a local fusion mode matching the real-time requirements; when the fusion mode is the global fusion mode, establishing correlations between the feature representations from all perspectives; when the fusion mode is the local fusion mode, establishing correlations between the feature representations from adjacent perspectives; and performing feature fusion based on the established correlations to obtain the fused global features.
[0010] In the implementation of the above scheme, by dynamically selecting the cross-view feature fusion mode according to the system performance requirements, an adaptive balance between computing resources and stitching quality is achieved. This enables the machine vision-based detection image stitching method to flexibly adapt to the technical requirements of different application scenarios, such as high-precision offline processing and real-time online processing. On the other hand, establishing the correlation between all view feature representations in the global fusion mode can fully utilize global context information to establish dependencies between arbitrary views, effectively handle parallax misalignment and exposure differences, and improve stitching accuracy and visual consistency in complex scenes. Furthermore, in the local fusion mode, only the correlation between adjacent view feature representations is established, reducing the computational complexity from global correlation to local neighborhood correlation. This significantly reduces computing power consumption while maintaining acceptable stitching quality, meeting the real-time constraints of high-density camera array deployment scenarios.
[0011] In one implementation of the first aspect, embedding viewpoint identification information during the encoding process includes: determining the azimuth position of each of the images to be stitched in the camera array; wherein, each camera in the camera array is horizontally and equally spaced; generating a relative azimuth position code based on the azimuth position; wherein, the relative azimuth position code is used to characterize the physical spatial topological relationship between each viewpoint; and fusing the relative azimuth position code with the feature encoding result of the corresponding viewpoint to obtain the feature representation of each viewpoint.
[0012] In the implementation of the above scheme, by encoding the physical layout of the camera array as a priori relative azimuth position code and embedding it into the feature representation, implicit modeling of the spatial topological relationship of each viewpoint is achieved under the condition of no explicit camera calibration, which improves the geometric accuracy and spatial consistency of uncalibrated image stitching. On the other hand, by determining the azimuth position of each image to be stitched in the camera array and generating the corresponding relative azimuth position code, each viewpoint feature is given a distinguishable spatial identity, which enhances the ability to distinguish inputs from different viewpoints and the ability to perceive orientation. Furthermore, by transforming the geometric constraint of horizontally equidistant layout into learnable coded information and integrating it into the feature representation, the model can use the physical prior knowledge of the regular array to guide feature alignment and fusion, reducing the dependence on large-scale labeled data and improving the convergence stability of model training.
[0013] In one implementation of the first aspect, the step of performing feature encoding on each of the proposed detection images to be stitched together to obtain feature representations from each viewpoint includes: extracting features at different spatial resolution levels to obtain multi-scale feature representations from each viewpoint; wherein the spatial resolution levels include an original resolution layer and at least one downsampled resolution layer; the step of performing cross-view feature fusion based on the correlation between the feature representations from different viewpoints to obtain fused global features includes: performing cross-view feature fusion based on the correlation between the multi-scale feature representations from different viewpoints to obtain fused global features.
[0014] In the implementation of the above scheme, by extracting features at different spatial resolution levels and performing cross-scale fusion, a comprehensive integration of multi-scale visual information is achieved, improving the ability of panoramic images to preserve details in multi-depth-of-field scenes. On the other hand, by extracting features at the original resolution level, high-frequency details and edge information of the image are preserved, avoiding detail loss caused by downsampling and improving the reconstruction accuracy of near objects and texture-rich areas. Furthermore, by extracting features at the downsampling resolution level, the receptive field is expanded and global contextual information is captured, enhancing the perception of distant objects and large-scale structural relationships, which is beneficial to improving the overall structural coherence of the panoramic image.
[0015] In one implementation of the first aspect, the method further includes: in response to a trigger signal, determining a target image from the detection images to be stitched from multiple perspectives; wherein the trigger signal includes an alarm signal generated by a perimeter security system or a panoramic video stitching request signal initiated by a user; the target image is the detection image to be stitched corresponding to the area indicated by the trigger signal; and performing a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step on the target image.
[0016] In the implementation of the above scheme, by responding to the trigger signal, stitching processing is only performed on the target image, realizing on-demand scheduling of computing resources and avoiding the ineffective computing power consumption caused by continuous stitching. This is beneficial to improving the engineering practicality of the above-mentioned machine vision-based detection image stitching method in large-scale camera array deployment scenarios. On the other hand, by receiving the alarm signal generated by the perimeter security system as the trigger condition, a panoramic image of the alarm area can be quickly generated when an intrusion event occurs, providing panoramic video verification support for security personnel, thereby improving the response speed and handling efficiency of security monitoring. Furthermore, by responding to the panoramic video stitching request signal initiated by the user, the stitching operation is only performed when the user needs to view the panoramic view of a specific area, reducing the amount of data processing and network transmission bandwidth occupation during unnecessary periods, and reducing the backend storage and computing load.
[0017] In one implementation of the first aspect, the method further includes: acquiring multi-view training images; synthesizing the multi-view training images into corresponding panoramic image labels to construct a training dataset; using the training dataset to train a model to obtain an image stitching model; wherein the image stitching model is configured to perform a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step.
[0018] In the implementation of the above scheme, by synthesizing multi-view training images into corresponding panoramic image labels to construct the training dataset, the data bottleneck problem of difficulty in obtaining real panoramic images and high labeling costs in large-scale ultra-wide-angle camera array scenarios is solved without relying on real panoramic image annotation, thus reducing the data threshold for model training. On the other hand, the image stitching model trained using synthetic data can learn the end-to-end mapping relationship from multi-view distorted images to panoramic images, enabling the model to have implicit modeling capabilities for ultra-wide-angle lens distortion characteristics, improving the model's generalization performance and stitching accuracy under uncalibrated conditions. Furthermore, by synthesizing panoramic image labels through geometric transformation or simulation rendering tools, training samples covering different scenes, different lighting conditions, and different camera layouts can be flexibly generated, enhancing the diversity of training data and the model's scene adaptability.
[0019] In one implementation of the first aspect, the method further includes: performing a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step locally on the detection host to obtain a panoramic image; encoding the panoramic image into a video stream and transmitting it to a backend server.
[0020] In the implementation of the above scheme, by performing feature encoding, cross-view feature fusion, and panoramic image generation steps locally on the detection host, the image stitching process is moved to the front-end device with edge computing power, realizing distributed processing of the computing load and reducing the centralized computing pressure and response latency of the back-end server. On the other hand, by encoding the generated panoramic image into a video stream instead of transmitting multiple original detection images to be stitched, the network transmission bandwidth occupied from the detection host to the back-end server is significantly reduced, improving the data transmission efficiency in large-scale camera array deployment scenarios. Furthermore, by using the edge computing power of the detection host to complete real-time stitching processing, the latency accumulation caused by the remote transmission of multiple ultra-wide-angle video streams is avoided, meeting the technical requirements of real-time panoramic video in security monitoring scenarios.
[0021] Secondly, embodiments of this application provide a machine vision-based detection image stitching device, comprising: The image acquisition module is used to acquire detection images to be stitched from multiple perspectives; The feature encoding module is used to encode the features of each of the proposed images to be stitched together, and to embed viewpoint identification information during the encoding process to obtain the feature representation of each viewpoint. The cross-view feature fusion module is used to perform cross-view feature fusion based on the correlation between the feature representations from different views, and obtain the fused global features; A decoding module is used to decode and generate a panoramic image from the global features based on a learnable query vector.
[0022] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a communication bus, wherein the processor and the memory communicate with each other through the communication bus; the memory stores computer program instructions that can be executed by the processor, and the computer program instructions are read and executed by the processor to perform the method provided in the first aspect or any possible implementation of the first aspect.
[0023] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the method provided in the first aspect or any possible implementation thereof.
[0024] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method provided by the first aspect or any possible implementation of the first aspect.
[0025] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims and drawings. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the perimeter security system provided in the embodiments of this application; Figure 2 A schematic flowchart illustrating the machine vision-based detection image stitching method provided in this application embodiment; Figure 3 A schematic diagram of the structure of the machine vision-based detection image stitching device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0028] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0029] Existing image stitching technologies generally employ a multi-stage, serial processing architecture based on local feature matching. This approach first extracts scale-invariant local feature points from images at various viewpoints using algorithms such as SIFT, SURF, or ORB. Then, feature descriptor matching is performed based on Euclidean distance or Hamming distance to establish correspondences between adjacent images. Subsequently, robust estimation methods such as RANSAC are used to calculate the homography matrix or camera pose parameters, and geometric transformations are used to map images at various viewpoints to a unified coordinate system. Finally, multi-band fusion or Poisson fusion algorithms are used to eliminate seams and exposure differences, generating a panoramic image.
[0030] However, in high-density node-based deployment scenarios for perimeter security, vibration detectors are typically deployed densely at intervals of 5 to 6 meters. A single detection host needs to process real-time video streams from multiple ultra-wide-angle cameras simultaneously (e.g., 10-in-1 stitching). Each frame of the image requires feature extraction, matching, and geometric estimation operations, resulting in computational complexity that increases linearly with the number of cameras. This makes it difficult to meet real-time requirements and the practical deployment needs of large-scale camera arrays. Furthermore, traditional local feature detection operators struggle to reliably extract discriminative feature points in areas with severe radial and tangential distortion at the edges of the ultra-wide-angle lens's field of view. Feature descriptors become distorted in these areas, and matching accuracy decreases. Simultaneously, geometric transformation estimation based on feature matching lacks sufficient constraints in edge regions, causing excessive loss or misalignment of edge image information during the fusion process, leading to degraded image stitching results.
[0031] In view of this, this application provides a machine vision-based detection image stitching method. This method replaces the traditional multi-stage geometric calibration and projection transformation process with an end-to-end deep learning architecture. Throughout the feature encoding, cross-viewpoint fusion, and panoramic generation processes, it does not rely on explicit calculations of camera intrinsic and extrinsic parameters and distortion models. It directly generates panoramic results from the original distorted image, avoiding edge information truncation and resampling distortion caused by explicit projection transformation, and achieving complete preservation of edge region information of ultra-wide-angle lenses. Furthermore, by embedding viewpoint identification information during feature encoding, the physical layout of the camera array is encoded a priori into a form that the model can perceive, enabling the model to understand the spatial relationships of each input image even without calibration. By establishing topological relationships and viewpoint identities, the accuracy and convergence stability of multi-view feature alignment are improved. Furthermore, cross-view feature fusion is achieved by establishing correlations between feature representations from different viewpoints. Global context modeling is used to replace the traditional local feature point matching mechanism, effectively handling parallax misalignment, dynamic occlusion, and exposure differences, thus improving the robustness and visual consistency of the stitching results in complex scenes. Additionally, a learnable query vector is used to directly decode and generate panoramic images from the fused global features. Adaptive feature sampling replaces the traditional back projection mapping and interpolation fusion process, avoiding interpolation holes, black border cropping, and ghosting issues, and achieving a direct mapping from the feature space to the image space.
[0032] like Figure 1As shown in the embodiments of this application, the machine vision-based detection image stitching method can be applied to a perimeter security system 100. This system is deployed in a perimeter security protection scenario with physical fences and achieves panoramic video monitoring and intrusion detection functions through a distributed multi-node collaborative architecture. The perimeter security system 100 can adopt a cloud-edge-device collaborative computing paradigm, consisting of a front-end perception layer, an edge processing layer, and a central analysis layer. The layers interact with each other with low latency through wired Ethernet or wireless communication networks, forming a complete technical closed loop from raw image acquisition and real-time panoramic stitching to intelligent analysis and decision-making.
[0033] Detectors 110 are deployed at equal intervals along the physical fence in a node-like manner, with the spacing between adjacent nodes typically set to 5 to 6 meters, forming a high-density monitoring array to eliminate physical blind spots. Detectors 110 integrate ultra-wide-angle miniature cameras with a horizontal field of view of no less than 180°, capable of covering a vast area inaccessible to traditional surveillance cameras. The optical centers of each camera are approximately on the same horizontal plane, with the main optical axis pointing horizontally and adjacent cameras maintaining a fixed angle, ensuring that images from adjacent viewpoints have more than 30% overlap in the horizontal direction, providing the necessary geometric and semantic basis for subsequent image stitching. Detectors 110 continuously acquire ultra-wide-angle raw images of the monitored area at a fixed frame rate (e.g., 25fps or 30fps) and transmit multiple video streams in real time to the nearest detection host 120 via an internal bus (e.g., RS485, CAN bus) or short-range wireless transmission methods (e.g., Wi-Fi, ZigBee), achieving local aggregation of front-end sensing data.
[0034] The detection host 120, as a field computing unit with edge computing capabilities, is typically deployed at the center of the detector cluster or in a nearby data center environment, primarily undertaking real-time processing tasks for multiple video streams. Its built-in heterogeneous computing hardware (such as GPU, NPU, or FPGA) loads an end-to-end image stitching model, performs feature encoding, cross-view feature fusion, and panoramic decoding generation operations, and synthesizes ultra-wide-angle images from multiple detectors 110 into a single seamless panoramic image or panoramic video sequence in real time.
[0035] Backend server 130, serving as the centralized management and analysis center of perimeter security system 100, is deployed in the monitoring center's computer room or cloud computing platform, receiving panoramic video streams from various detection hosts 120. Backend server 130 can perform advanced video analysis tasks based on the received panoramic images, including but not limited to intrusion behavior recognition, target trajectory tracking, abnormal event detection, and video verification analysis based on deep learning algorithms; it also handles system configuration management, user access control, data storage and archiving, and visualization. Since the panoramic images have already been fused and generated at the detection host 120, backend server 130 does not need to process the original multi-channel distorted images or perform complex cross-view stitching operations. It only needs to focus on intelligent analysis based on a single panoramic video stream, thereby significantly reducing the computational load, storage pressure, and network bandwidth requirements of the central node, and improving the overall system response efficiency, architectural scalability, and economic feasibility in large-scale node deployment scenarios.
[0036] like Figure 2 As shown in the figure, this application provides a machine vision-based detection image stitching method, which includes: Step S210: Obtain detection images to be stitched from multiple perspectives.
[0037] In perimeter security scenarios, detectors 110 are deployed at equal intervals along the physical fence in a node-like manner. Each detector 110 integrates an ultra-wide-angle miniature camera. Images from multiple perspectives to be stitched together can be acquired by the ultra-wide-angle miniature cameras mounted on the detectors 110. The installation orientation of each ultra-wide-angle miniature camera satisfies consistency constraints: the optical centers of all cameras are approximately on the same horizontal plane, the principal optical axes point horizontally, and the principal optical axes of adjacent cameras maintain a fixed angle, ensuring that images from adjacent perspectives have a certain degree of overlap in the horizontal direction. This overlap provides the necessary geometric and semantic association information for subsequent image stitching, enabling images from different perspectives to be accurately aligned and fused.
[0038] Under normal operating conditions, detector 110 continuously acquires video streams of the monitored area. When the detection host 120 receives an intrusion detection alarm signal or a panoramic viewing request initiated by a user, it triggers an image stitching process, acquiring the current frame of the detection images to be stitched from the corresponding multiple detectors 110. These images retain the inherent radial distortion characteristics of ultra-wide-angle lenses and are directly used as input to the end-to-end image stitching model without explicit distortion correction or preprocessing. The detection host 120 then performs feature encoding, cross-view fusion, and panoramic decoding generation operations.
[0039] Step S220: Perform feature encoding on each detection image to be stitched together, and embed viewpoint identification information during the encoding process to obtain the feature representation of each viewpoint.
[0040] The aforementioned feature encoding is a nonlinear transformation process that maps high-dimensional image data in the original pixel space to a low-dimensional semantic feature space. It aims to extract high-level semantic representations robust to illumination changes, local deformations, and noise interference, while simultaneously compressing data dimensionality to reduce the complexity of subsequent cross-viewpoint correlation calculations. In the end-to-end image stitching architecture, feature encoding, as the first-layer processing module, plays a crucial role in converting ultra-wide-angle distorted images from various viewpoints into a unified feature space representation, providing a computable feature foundation for subsequent cross-viewpoint geometric alignment and semantic fusion.
[0041] Feature encoding can be implemented using a Vision Transformer-based architecture, a deep learning model based on the Transformer architecture. The Vision Transformer uniformly segments the input image into a sequence of non-overlapping two-dimensional image patches. Each patch, after being flattened, is mapped to a high-dimensional feature vector through a learnable linear projection matrix, forming an initial visual token representation. The linear embedding layer compresses the texture and structural information of the local pixel neighborhood into a dense vector, preserving the spatial locality prior of the image, while ensuring consistency and parameter efficiency in feature extraction across different viewpoints through a weight-sharing mechanism.
[0042] The Vision Transformer comprises an image serialization module, a positional encoding layer, and stacked Transformer encoder layers. The image serialization module divides the input image into fixed-size two-dimensional image patches through uniform segmentation, flattens each patch into a one-dimensional vector, and maps it to a high-dimensional feature space via a learnable linear projection matrix, forming a visual token sequence. The positional encoding layer adds two-dimensional spatial location information to each visual token using learnable embedding vectors or sinusoidal positional encoding, compensating for the spatial structure loss caused by serialization. The core processing unit of the Vision Transformer includes multiple stacked encoder layers with identical structures. Each layer contains a multi-head self-attention module, a multilayer perceptron feedforward network, and corresponding layer normalization and residual connection components. During operation, the visual token sequence is first added element-wise with the positional encoding to form an initial feature representation with spatial prior. This initial feature representation is input to the first encoder layer, where the multi-head self-attention module linearly projects the features into three matrices: query, key, and value. Attention weights are obtained by calculating the scaled dot product of the query and key, and then the value matrix is weighted and aggregated to achieve direct association and global dependency modeling between features of any two image patches. Subsequently, the multilayer perceptron feedforward network performs position-wise nonlinear feature transformation on the self-attention output, expanding the feature representation capability through two layers of linear mapping and activation functions. The output of each sublayer (self-attention module and feedforward network) is added to the input features through residual connections, and the data distribution is stabilized by layer normalization to form the final feature representation of that layer, which is then passed to the next layer. Through this hierarchical stacking and global attention computation, the Vision Transformer abstracts hierarchical features from local texture details to global semantic structure layer by layer, ultimately outputting a high-dimensional feature representation sequence for each image patch, which serves as the input representation for downstream cross-view fusion tasks.
[0043] Besides the Vision Transformer, feature encoding can also be implemented using a convolutional neural network-based encoding architecture. For example, using ResNet or DenseNet as the backbone network, a hierarchical feature pyramid is constructed by stacking convolutional layers, batch normalization layers, and nonlinear activation functions. Downsampling operations are used to gradually expand the receptive field and extract multi-scale semantic features. Finally, global average pooling or flattening operations are used to compress the feature map into a fixed-dimensional feature vector. This process also supports end-to-end gradient backpropagation and parameter optimization. Another feasible implementation is to first use a lightweight convolutional neural network (such as MobileNet or EfficientNet) to extract local texture and edge features of the image, and then use a Transformer encoder to model the long-range dependencies between features. By combining the local perception of convolution with the global modeling of attention, the computational complexity is reduced while maintaining end-to-end training characteristics.
[0044] The aforementioned viewpoint identification information is a learnable embedding vector used to explicitly represent the viewpoint identity and spatial topological relationships of each input image in a multi-camera array. Viewpoint identification information is typically fused with the visual feature representation of the image in the form of a high-dimensional vector, assigning a unique identity and relative orientation attribute to the feature representation of each viewpoint. This allows the feature vector, which originally only contained pixel semantic information, to carry prior knowledge of the physical spatial layout, forming a structured multi-viewpoint feature representation.
[0045] In the above scheme, the operation of embedding viewpoint identification information during the encoding process aims to enable the end-to-end model to perceive the relative positional relationship and array topological order of each input image in physical space even without explicit camera intrinsic and extrinsic parameter calibration, thereby establishing cross-viewpoint geometric associations and correspondences at the feature level.
[0046] In this application embodiment, viewpoint identification information can be obtained through at least one of the following multiple implementation methods: The first implementation method: obtains the position through relative azimuth position encoding; Optionally, embedding viewpoint identification information during the encoding process includes: determining the azimuth position of each image to be stitched in the camera array; wherein, each camera in the camera array is horizontally and equally spaced; generating a relative azimuth position code based on the azimuth position; wherein, the relative azimuth position code is used to characterize the physical spatial topological relationship between each viewpoint; and fusing the relative azimuth position code with the feature encoding result of the corresponding viewpoint to obtain the feature representation of each viewpoint.
[0047] Relative azimuth position encoding is a position embedding vector generated based on the geometric layout of a camera array. It is used to explicitly represent the horizontal deflection angle of the principal optical axis of each camera relative to a reference direction and their interrelationships. The azimuth position can be defined as the angle between the optical axis of each camera in the horizontal plane and a preset reference direction (such as the array's starting direction or true north). In a horizontally evenly spaced camera array, the azimuth difference between adjacent cameras remains constant, forming a regular angular sampling sequence. Relative azimuth position encoding then maps this continuous angle value to a discrete or continuous representation in a high-dimensional vector space using sine / cosine functions or a learnable embedding matrix, enabling the encoded vector to carry relative azimuth and distance information between viewpoints.
[0048] The encoding process mainly includes: (1) Determining the azimuth position of each image to be stitched according to the calibration parameters or design drawings of the camera array. This position information reflects the spatial ownership of the current viewpoint in the array; (2) Generating a relative azimuth position code based on the azimuth position. The generation method can adopt a deterministic mapping based on sine position coding, decompose the azimuth into sine and cosine components and extend it to a high-dimensional space, or adopt a parametric mapping based on learnable embedding, and obtain the optimal code through training optimization; (3) Adding the relative azimuth position code to the feature representation obtained by visual feature extraction of the corresponding viewpoint element by element or channel stitching and fusion to form a structured feature representation with physical space prior.
[0049] In the above scheme, since the azimuth interval between adjacent cameras is fixed, the relative azimuth position encoding enables the model to perceive the arrangement order and adjacency relationship of each viewpoint in physical space through the similarity or relative position relationship of the encoded vectors, even without explicit extrinsic parameter calibration. Thus, in the subsequent cross-view feature fusion stage, strong feature associations between physically adjacent viewpoints are established first, and the azimuth differences implied by the encoding are used to assist geometric alignment, thereby achieving calibration-free image stitching based on physical layout priors.
[0050] The second implementation method: adopting a learnable perspective identity embedding vector method; In this embodiment, a unique learnable embedding vector can be pre-assigned to each detector 110. This vector serves as a model parameter and is optimized and updated together with the network weights during training. In the feature encoding stage, the viewpoint-specific embedding vector is added or concatenated element-wise with the visual feature representation of the image patch to form a feature representation with viewpoint identity.
[0051] The third implementation method: obtaining the data based on the index encoding method of the camera array; In this embodiment, a unique integer index can be assigned to each viewpoint based on the deployment order or node number of each detector in the array. This index is mapped to a high-dimensional dense vector through an embedding layer, or a one-hot encoded vector can be directly used as the viewpoint identifier. During feature encoding, the index encoding is fused with visual features, enabling the model to distinguish different viewpoints and perceive adjacent relationships based on the array index.
[0052] In addition to viewpoint identification information, positional identification information can also be embedded during the encoding process. Positional identification information is an embedding vector used to encode the spatial coordinates of visual tokens in the original two-dimensional image plane. Its role is to add spatial positional attributes to the features of image patches after serialization. In the Vision Transformer architecture, the input image is uniformly divided into regular two-dimensional grid-like image patches. After each image patch is linearly projected into a visual token, the positional identification information, through learnable embedding parameters or deterministic mathematical functions, encodes the row and column indices or two-dimensional coordinates of the image patch into a high-dimensional vector representation, and then fuses it with the feature vector of the visual token. It can be understood that because the Vision Transformer segments and flattens the image into a one-dimensional sequence, the inherent spatial adjacency relationships and geometric layout information of the original image are lost, and the model cannot distinguish visual tokens from the upper left and lower right regions of the image. By introducing positional identification information, the model can reconstruct the relative distances, row and column positions, and spatial topology between image patches at the feature level, thereby introducing spatial prior constraints in the self-attention computation. This ensures that the feature extraction process accurately preserves the geometric consistency and spatial resolution of the original image, improving the processing accuracy for spatially sensitive tasks.
[0053] The aforementioned location identification information can be achieved using a deterministic encoding method based on sine and cosine functions. This method uses sine and cosine components of different frequencies to periodically map the two-dimensional coordinates of image patches, generating encoded vectors with relative position awareness. Alternatively, a parameterized method based on learnable embedding matrices can be used. This maps discrete grid coordinates into dense vectors optimized through end-to-end training, allowing the location encoding to adapt to the spatial distribution characteristics of a specific dataset. Both methods fuse with visual features through element-wise addition or channel concatenation operations, forming a composite feature representation that combines semantic information and spatial coordinates.
[0054] It is understandable that location identification information and viewpoint identification information can complement and synergize in terms of functionality: the former focuses on spatial topological encoding within a single image, representing the relative geometric relationships between image patches, enabling the model to possess refined local spatial perception capabilities; the latter focuses on cross-camera viewpoint identification, representing the viewpoint affiliation of each image within the physical array. When both are embedded together, the model simultaneously possesses the ability to perceive spatial location within an image and the ability to correlate geometrically across viewpoints, providing a complete prior information foundation for subsequent cross-viewpoint feature alignment and fusion, effectively improving the spatial accuracy and structural consistency of multi-viewpoint image stitching under uncalibrated conditions.
[0055] Optionally, the above-mentioned method for obtaining feature representations includes: extracting features at different spatial resolution levels to obtain multi-scale feature representations from each perspective; wherein the spatial resolution level includes an original resolution layer and at least one downsampled resolution layer; and performing cross-perspective feature fusion based on the correlation between feature representations from different perspectives to obtain fused global features, including: performing cross-perspective feature fusion based on the correlation between multi-scale feature representations from different perspectives to obtain fused global features.
[0056] The aforementioned multi-scale feature representation refers to the feature sets extracted from the input image and its downsampled image, forming a feature pyramid structure spanning different spatial resolution levels. The spatial resolution levels include the original resolution layer, which preserves pixel-level details, and the downsampled resolution layer, which reduces spatial dimensions and expands the receptive field through pooling or stride convolution operations. The original resolution layer is rich in high-frequency edge information and fine textures, while the downsampled resolution layer contains global semantic context and large-scale geometric structures. These two layers complement each other in information representation, jointly providing a multi-granular feature foundation for cross-viewpoint alignment. Multi-scale feature extraction can be achieved through various network architectures. One optional implementation uses a feature pyramid network structure, fusing deep low-resolution semantic features with shallow high-resolution positional features via a top-down path, and utilizing lateral connections to ensure effective transmission of detailed information during upsampling. Another optional implementation uses an encoder-decoder architecture, where the encoder generates multi-level feature maps through continuous downsampling, and the decoder receives encoder features from the same level through skip connections, achieving direct fusion of multi-scale information. Furthermore, a parallel multi-branch structure can also be used, where each branch independently processes inputs at different resolutions and extracts feature representations at the corresponding scale.
[0057] Cross-view fusion based on multi-scale feature representation can be achieved using one of the following strategies: One strategy is to establish the correlation between features from different perspectives at each spatial resolution level and perform fusion to generate multi-scale fused features before hierarchical aggregation; another strategy is to adopt a coarse-to-fine hierarchical fusion strategy, specifically: first, establish cross-view correlations at the lower resolution level to obtain a coarse geometric alignment relationship, and then pass this alignment information to the higher resolution level to guide fine feature fusion; yet another strategy is to stitch or weightedly fuse features from different scales and then uniformly establish cross-view correlations to achieve cross-scale information interaction.
[0058] In the above scheme, the high-frequency details provided by the original resolution layer help to accurately locate corresponding points in overlapping areas and preserve the sharpness of edge structures, avoiding detail loss and seam blurring caused by downsampling. The downsampling resolution layer captures a wide range of geometric structures and semantic context by expanding the receptive field, which helps to handle overall alignment and exposure difference correction in scenes with large parallax. By establishing the relationship between viewpoints at different resolution levels, the model can simultaneously utilize local fine features and global structural priors to achieve coarse constraints on large displacements and fine adjustments to small deviations, thereby improving the overall quality and spatial consistency of panoramic image generation.
[0059] Step S230: Based on the correlation between feature representations from different perspectives, perform cross-perspective feature fusion to obtain the fused global features.
[0060] The aforementioned cross-perspective feature fusion is a process of integrating feature representations from multiple independent perspectives into a unified global feature representation. Its aim is to establish semantic and geometric connections between different perspectives and eliminate the spatial isolation of single-perspective features. These connections refer to the correspondence, similarity, or spatial dependency between feature representations from different perspectives. They characterize the degree of matching between corresponding points in overlapping areas of each perspective in physical space and their semantic content. By calculating the similarity between feature vectors or establishing constraints using geometric priors, a foundation for cross-perspective feature interaction is formed.
[0061] One possible implementation of cross-perspective feature fusion is to employ a cross-attention mechanism. This mechanism uses the feature representation of one perspective as a query vector and the feature representations of other perspectives as key and value vectors. By calculating the similarity weights between the query and the keys, the value vectors are weighted and aggregated, enabling adaptive sampling and fusion of features from other perspectives from the current perspective. This attention mechanism supports multi-head parallel computation, can simultaneously capture the relationships between different subspaces, and allows bidirectional or multidirectional information flow, enabling features from different perspectives to update and enhance each other during interaction.
[0062] Another alternative implementation for cross-view feature fusion is a feature aggregation approach based on graph neural networks. In this implementation, the feature representation of each viewpoint can be considered as a node in a graph structure. Edge connections between nodes are constructed based on physical layout or feature similarity, forming a cross-view graph topology. Through graph convolution or graph attention networks, each node aggregates the feature information of its neighboring nodes and iteratively updates its own feature representation, enabling the features of related views to be smoothly fused in the local neighborhood, ultimately generating a global feature representation with spatial consistency.
[0063] By establishing and fusing the relationships between feature representations from different perspectives, the model can implicitly learn the correspondence between the geometric layout of the camera array and the perspective at the feature level, achieving geometric alignment of multi-view features without explicit parameter calibration. Through cross-view information interaction, features in overlapping areas are mutually validated and enhanced, while features in non-overlapping areas receive reasonable semantic completion through contextual association. The resulting global feature representation carries complete scene geometric and semantic information, providing a unified feature foundation for subsequent query vector-based panoramic decoding and ensuring the consistency and coherence of the generated panoramic image in terms of content and structure.
[0064] Optionally, step S230 includes: determining a fusion mode for cross-view feature fusion based on system performance requirements; wherein, system performance requirements include high precision requirements and real-time requirements; the fusion mode includes a global fusion mode matching the high precision requirements and a local fusion mode matching the real-time requirements; when the fusion mode is a global fusion mode, establishing the correlation between the feature representations of all views; when the fusion mode is a local fusion mode, establishing the correlation between the feature representations of adjacent views; and performing feature fusion based on the established correlation to obtain the fused global features.
[0065] The aforementioned fusion mode refers to the abstract representation of the scope and strategy for establishing feature relationships during cross-view feature fusion, which determines the breadth and computational complexity of information interaction between different viewpoints. System performance requirements are constraints guiding the selection of fusion modes, including two categories: high-precision requirements and real-time requirements. High-precision requirements correspond to offline processing or key area monitoring scenarios with strict requirements for stitching quality, while real-time requirements correspond to large-scale online video stream processing scenarios that are sensitive to processing latency. The global fusion mode establishes fully connected relationships between all available viewpoint feature representations, enabling direct information interaction between any two viewpoints; the local fusion mode, on the other hand, only establishes local relationships between physically adjacent viewpoint feature representations, limiting the interaction scope to neighboring viewpoint pairs with overlapping areas.
[0066] When determining the fusion mode, decisions can be made based on preset configuration parameters or dynamically monitored current load status. Specifically: when a command signal requiring high precision is received or sufficient computing resources are detected, the global fusion mode is activated, calculating the association weights of all view pairs and performing feature fusion; when a command signal requiring real-time performance is received or a decrease in the current processing frame rate is detected, the mode switches to local fusion, calculating only the association weights of adjacent view pairs. The above scheme can achieve global fusion through a multi-head self-attention mechanism, calculating the global attention weight matrix by using the feature representations of each view as a unified token sequence input. Furthermore, the above scheme can achieve local fusion through window attention or mask attention, shielding connections between non-adjacent views by retaining only the attention calculation paths between adjacent views.
[0067] In the above scheme, the fusion mode is dynamically selected based on system performance requirements, and a corresponding range of correlations is established, achieving an adaptive balance between computing resources and stitching quality. The global fusion mode, by fully utilizing the global context information of the camera array, can establish long-range dependencies between arbitrary viewpoints, effectively handling challenging scenarios with large parallax, complex occlusion, and significant exposure differences, improving the geometric accuracy and visual consistency of the stitching results. The local fusion mode, by limiting the scope of correlation calculations, reduces computational complexity from global calculations related to the square of the number of viewpoints to local calculations linearly related to the number of viewpoints. While maintaining acceptable stitching quality, it significantly reduces computing power consumption and processing latency, meeting the real-time constraints of high-density camera array deployment scenarios. The dynamic switching between the two modes allows the above machine vision-based detection image stitching method to flexibly adapt to the resource conditions of actual application scenarios, achieving both high precision and high efficiency.
[0068] Step S240: Generate a panoramic image from global features based on the learnable query vector.
[0069] The aforementioned query vector is a learnable embedding vector that serves as a structured query unit in the decoding stage. It adaptively aggregates information from the fused global features to generate pixel representations of specific spatial locations or regions in the panoramic image. Unlike the deterministic mapping based on fixed-grid upsampling in traditional decoders, the query vector establishes a dynamic relationship between the output spatial location and the input feature space through parametric learning. This enables adaptive sampling and selective aggregation of global features, thereby transforming high-dimensional abstract feature representations into regular gridded image pixel outputs.
[0070] During decoding, each query vector is explicitly configured as a spatial location in the corresponding panoramic image, responsible for generating the pixel value or local region feature at that location. The decoder establishes a correlation strength measure between the query vector and global features by calculating the interaction weights between them. Then, based on these weights, it performs weighted summation or attention aggregation on the global features to extract the feature information most relevant to the current output location. Finally, it maps the aggregated feature vector to pixel values through linear projection or a lightweight network, forming the corresponding location output of the panoramic image. This computational process supports large-scale parallel computing, and the decoding operations of each query vector are independent of each other, enabling the simultaneous generation of all pixels in the panoramic image, thereby improving decoding efficiency.
[0071] In the initial state, the query vector can be randomly initialized or sampled based on the prior distribution of the output spatial grid. During model training, the pixel-level reconstruction loss or perceptual loss between the generated panoramic image and the target panoramic image is defined, the gradient of the loss function with respect to the query vector parameters is calculated, and the value of the query vector is updated based on the gradient descent algorithm, so that it gradually converges to the optimal state that can accurately capture the visual semantic information required at each position of the panoramic image.
[0072] The above scheme employs learnable query vectors for panoramic decoding, establishing a direct generation path from the feature space to the image space. This eliminates the post-processing steps of geometric transformation-based back projection mapping, interpolation filling, and multi-stage fusion found in traditional methods. By automatically learning a mapping strategy from global features to panoramic pixels through data-driven learning, this scheme avoids the problems of edge information truncation and interpolation holes caused by explicit projection, eliminating the need for explicit camera calibration parameters and distortion models. Furthermore, the adaptive aggregation characteristics of the query vectors enable the model to dynamically adjust the sampling area and weight allocation based on the current scene content, prioritizing key feature information semantically related to the output location. This improves the flexibility, content adaptability, and spatial consistency of panoramic image generation, ensuring the integrity of the generated results in terms of geometric structure and visual quality.
[0073] Optionally, step S240 includes: dynamically predicting the sampling position and offset of each learnable query vector in the global feature space; sampling and aggregating feature information from the global features based on the sampling position and offset to generate corresponding panoramic image pixels; wherein each query vector is configured to generate panoramic image pixels at the corresponding spatial position in the panoramic image.
[0074] The aforementioned sampling position and offset are parameter pairs used to precisely define the information extraction coordinates in the global feature space. The sampling position serves as a reference point, typically corresponding to the initial projected coordinates of the query vector in the global feature map. The offset is a learnable or predictable displacement vector generated relative to this reference point, representing the dynamic adjustment from the initial reference point to the actual information-rich region. The sampling position and offset are superimposed to form the final sampling coordinates, enabling the query vector to overcome the limitations of fixed grid sampling and freely locate itself to the most relevant feature region in the global feature space.
[0075] In the above scheme, a lightweight prediction network can dynamically predict the sampling position and offset. The prediction network takes a learnable query vector as input and calculates the output offset through a linear transform layer or a small multilayer perceptron. The prediction process is jointly trained end-to-end with the main network of the panoramic decoding. In another implementation, the offset prediction can be combined with the panoramic image output position encoding corresponding to the query vector. A cross-attention mechanism is used to calculate the relevance of each global feature position to the current query, thereby deriving the optimal sampling offset distribution. The numerical range of the offset can be limited by a learnable temperature parameter or hard constraints to ensure that the sampling points are always within the effective spatial domain of the global features.
[0076] An optional implementation method for sampling and aggregating feature information from global features based on sampling location and offset includes: extracting feature vectors at predicted floating-point coordinates using bilinear or bicubic interpolation; then, weighting and summing the features of multiple sampling points using an attention weighting mechanism to form a comprehensive feature representation of the query location. This comprehensive feature representation is mapped to the pixel values of the corresponding spatial locations in the panoramic image via a feedforward network, completing the transformation from abstract features to visual output. Each query vector maintains a fixed one-to-one correspondence with the spatial location in the panoramic image, ensuring a clear spatial responsibility division during the decoding process and avoiding pixel position confusion and misalignment.
[0077] In the above scheme, a mechanism that dynamically predicts sampling positions and offsets endows the query vector content with adaptive spatial awareness. This allows the query vector to actively search for and aggregate the most discriminative information regions in the global features based on the geometric complexity and semantic distribution of the current scene, rather than passively accepting the average features within a fixed neighborhood. This sparse sampling strategy significantly reduces invalid computation and noise interference in the global feature space, improving the accuracy and efficiency of feature aggregation. Simultaneously, by dynamically adjusting the sampling coordinates, the model can effectively handle nonlinear distortion regions at the edges of ultra-wide-angle lenses and overlapping regions with large parallax, avoiding detail loss and geometric distortion caused by fixed grid sampling. This ensures that the generated panoramic image pixels achieve a higher level of accuracy in content and structural integrity.
[0078] Optionally, the above-mentioned machine vision-based detection image stitching method further includes: in response to a trigger signal, determining a target image from multiple viewpoint detection images to be stitched; wherein the trigger signal includes an alarm signal generated by a perimeter security system or a panoramic video stitching request signal initiated by a user; the target image is the detection image to be stitched corresponding to the area indicated by the trigger signal; and performing a feature encoding step, a cross-viewpoint feature fusion step, and a panoramic image generation step on the target image.
[0079] The aforementioned trigger signal refers to the asynchronous event-driven instruction received by the detection host 120. The trigger signal can serve as the logical starting condition for initiating the panoramic stitching process. Based on its source, the trigger signal includes alarm signals automatically generated by the perimeter security system 100 when an intrusion event is detected, and panoramic video stitching request signals actively initiated by the user through the human-machine interface when reviewing video of a specific area. Alarm signals are typically transmitted to the detection host 120 via the system's internal bus or network protocol in the form of electrical signals or data packets, carrying identification information of the trigger area or detector node. Panoramic video stitching request signals can be generated by user operation instructions on the backend server 130 or client terminal, carrying the geographical area or camera array range to be viewed panoramically.
[0080] Upon receiving a trigger signal, the region identifier or coordinate information carried in the signal can be parsed, and a pre-configured detector array topology mapping table can be queried to determine the target image corresponding to the region indicated by the trigger signal from multiple perspectives of the detector images to be stitched together. The target image specifically refers to the real-time frame image captured at the current moment by an ultra-wide-angle miniature camera on a specific detector node covered by the trigger region. In an optional implementation, a correspondence between detector node IDs and physical coordinate ranges or logical area codes can be pre-established in a database or configuration file. When the region identifier in the alarm signal or user request matches the coverage area of a specific node or a cluster of consecutive nodes, the image stream from the corresponding camera can be marked as the target image, while the image streams from nodes outside the trigger region are filtered out to prevent them from entering the subsequent stitching processing pipeline.
[0081] For the target image, a pre-trained image stitching model can be invoked to sequentially execute feature encoding, cross-view feature fusion, and panoramic image generation steps. This processing flow remains logically consistent with processing all detected images, but differs in the input data range; the model only performs computational graph unfolding and parameter inference for a specific subset of viewpoint images. At the hardware implementation level, GPU or NPU computing resources can be dynamically allocated based on the number of target images, loading model weights for the corresponding subset or activating specific computational paths, thereby avoiding the computational overhead and memory consumption associated with processing the entire camera array. In scenarios with high-density vibration detector deployment, when only a local area triggers an alarm, 2-in-1, 3-in-1, or even N-in-1 local panoramic videos can be generated based on the alarm area coverage, rather than being limited to full array stitching.
[0082] The above solution avoids the massive resource consumption associated with continuous full-camera panoramic stitching in high-density node deployment scenarios by responding to trigger signals and processing target images of specific areas. This response mechanism allows the detection host to concentrate its limited edge computing power on the area where the intrusion event occurred or the area of user interest, generating panoramic video with low latency to support real-time alarm verification and security decision-making. Simultaneously, it reduces the transmission of full panoramic video streams in non-triggered states, lowers the load on backend servers and network links, and ensures priority service quality for data streams in key areas under limited bandwidth conditions, thereby improving the engineering applicability and economic feasibility of large-scale perimeter security systems.
[0083] Optionally, the above-mentioned machine vision-based detection image stitching method further includes: acquiring multi-view training images; synthesizing the multi-view training images into corresponding panoramic image labels to construct a training dataset; using the training dataset to train the model and obtain an image stitching model; wherein the image stitching model is configured to perform a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step.
[0084] The image stitching model described above adopts an end-to-end three-layer cascaded architecture of encoder-feature fusion module-decoder, which enables direct mapping from multiple ultra-wide-angle distorted images to a single seamless panoramic image. The image stitching model abandons the explicit geometric calibration, projection transformation, and post-processing fusion modules in traditional multi-stage processing flows. Instead, it achieves joint optimization of feature extraction, cross-viewpoint alignment, and image generation through neural networks, supporting end-to-end training based on gradient descent. The encoder can employ a parameter-shared Vision Transformer encoder, responsible for mapping the input images from each viewpoint from pixel space to a high-dimensional semantic feature space. The main processing flow includes: first, uniformly dividing the image into a sequence of two-dimensional image blocks; converting the image blocks into visual tokens through a linear embedding layer; and superimposing position encoding and viewpoint identifier encoding to form a structured input representation; subsequently, through a multi-layer stacked self-attention mechanism and a multi-layer perceptron feedforward network, extracting hierarchical features from local texture to global semantics layer by layer, outputting a high-dimensional feature representation sequence for each viewpoint. The feature fusion module, deployed after the encoder, is responsible for establishing the relationships between feature representations from different perspectives and performing information aggregation. This module can be implemented based on a cross-attention mechanism or a graph neural network. By calculating the similarity weights between features from different perspectives, it establishes cross-perspective association edges, enabling features from different perspectives to adaptively exchange information based on prior physical layout and semantic correspondence. In global fusion mode, it establishes fully connected associations between all perspectives; in local fusion mode, it establishes only local associations between adjacent perspectives. It generates global fusion features carrying complete scene information through weighted aggregation or message passing mechanisms. The image stitching model also includes a set of learnable panoramic query vectors and a decoder based on deformable attention. Each query vector is configured to generate pixels at a specific spatial location in the panoramic image. By dynamically predicting the sampling position and offset in the global feature space, it adaptively aggregates the most relevant visual information from the fusion features. The aggregated features are mapped to pixel values via a feedforward network, and the results of all query positions are output in parallel to form a regular gridded panoramic image.
[0085] The aforementioned multi-view training images refer to sample data collected from cameras at various viewpoints within a camera array, used for model learning. These typically include multiple ultra-wide-angle raw images, with their physical layout satisfying constraints such as approximately coplanar optical centers, horizontally spaced principal optical axes, and at least 30% overlap between adjacent lenses. Panoramic image labels refer to the ground truth panoramic images corresponding to the multi-view training images, serving as the regression target for supervised learning and representing the ideal panoramic result that should be obtained after stitching together these multi-view images. The training dataset consists of paired sets of multiple multi-view training images and their corresponding panoramic image labels, used for optimizing end-to-end network parameters.
[0086] In scenarios lacking real panoramic image labels, multi-view training images can be synthesized into corresponding panoramic image labels using geometric transformations or simulation rendering tools. Specifically: based on known camera array geometric parameters or approximate calibration results, perspective transformations or spherical projections are used to map images from each viewpoint to a unified cylindrical or spherical coordinate system, and image fusion algorithms are used to generate synthesized panoramic images. Alternatively, 3D scene reconstruction and virtual camera rendering techniques can be employed to construct a virtual camera array in a simulation environment that is geometrically consistent with the real scene, and distortion-free panoramic image labels can be directly generated through ray tracing or rasterization rendering. During the model training phase, an end-to-end supervised learning framework is used. The network, which includes feature encoding, cross-view fusion, and query decoding, is iteratively optimized using the constructed training dataset. The network parameters are updated by minimizing the reconstruction loss between the generated panoramic image and the panoramic image label until the model converges to obtain the image stitching model.
[0087] The aforementioned scheme constructs a training dataset through synthesis and implements end-to-end model training. Even in the absence of large-scale real-world labeled data, it enables the model to learn the direct mapping relationship from multi-view distorted images to panoramic images. Furthermore, this scheme avoids the high cost and difficulty of acquiring real panoramic images and manual annotation. Synthetic data generated using geometric transformations or simulation tools provides ample supervision signals, allowing the model to implicitly learn the geometric layout of the camera array, the nonlinear distortion characteristics of the ultra-wide-angle lens, and cross-view correspondences during training. Therefore, in practical deployment, it eliminates the need for explicit camera calibration parameters and manually designed geometric transformation rules, directly achieving high-quality image stitching based on the learned parametric mapping. This improves the data availability and model generalization ability of the aforementioned machine vision-based image stitching method in engineering practice.
[0088] Optionally, the above-mentioned machine vision-based detection image stitching method further includes: performing a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step locally on the detection host to obtain a panoramic image; encoding the panoramic image into a video stream and transmitting it to a backend server.
[0089] The aforementioned detection host 120 is an edge computing unit deployed at the monitoring site, possessing independent graphics processing or neural network inference capabilities. It is used to perform real-time image processing and analysis tasks near the data source. In the perimeter security system 100, the detection host 120 and multiple detector nodes 110 form a local cluster, directly receiving raw video streams from ultra-wide-angle cameras and undertaking computationally intensive operations such as image stitching, rather than transmitting multiple raw video streams to a remote central node for processing. The local execution of feature encoding, cross-view feature fusion, and panoramic image generation steps means that the complete end-to-end stitching process is completed in a closed loop on the local hardware resources of the detection host, generating single-channel panoramic image data.
[0090] The probe host 120 can utilize integrated heterogeneous computing units such as GPUs, NPUs, or FPGAs to load pre-trained image stitching model weights and perform parallel inference computation on the input multi-channel video frames. The feature encoding step achieves rapid embedding of visual tokens and efficient computation of the Transformer layer through hardware-accelerated matrix operations; the cross-view feature fusion step utilizes the global feature cache in the device's memory to establish cross-view attention associations and perform feature aggregation; the panoramic image generation step outputs pixel-level panoramic image frames through a query decoder. During the above processing, the probe host 120 does not need to interact with the backend server 130 via intermediate data; all computational dependencies are completed in local memory and storage. After generating the panoramic image, the probe host 120 encodes the single-channel panoramic image frame sequence into a standard video stream format. It can use compression encoding algorithms such as H.264 or H.265 to perform lossy compression on the panoramic video, reducing data size to adapt to network transmission bandwidth constraints; then, it pushes the encoded video stream to the backend server via real-time transmission protocols such as RTSP or RTMP. Compared to transmitting multiple raw ultra-wide-angle video streams, the data volume of a single panoramic video stream is reduced, and the need for the backend server 130 to perform stitching calculations is eliminated.
[0091] The above solution performs image stitching processing locally on the detection host 120 and only transmits the panoramic results, achieving a reasonable distribution of computational load and efficient utilization of network bandwidth. Local processing avoids long-distance transmission delays and bandwidth consumption of multiple high-resolution video streams. In addition, the backend server 130 receives the fused panoramic image, eliminating the need to configure a high-performance computing cluster to perform complex cross-view registration and fusion operations. It can focus solely on intrusion detection analysis or storage management based on the panoramic image, thereby reducing the hardware cost and architectural complexity of the perimeter security system 100 and improving its real-time response capability and engineering scalability.
[0092] like Figure 3 As shown, based on the same inventive concept, this application also provides a machine vision-based detection image stitching device 300, which includes: Image acquisition module 310 is used to acquire detection images to be stitched from multiple perspectives; The feature encoding module 320 is used to encode the features of each detection image to be stitched together, and to embed viewpoint identification information during the encoding process to obtain the feature representation of each viewpoint; The cross-view feature fusion module 330 is used to perform cross-view feature fusion based on the correlation between feature representations from different perspectives to obtain the fused global features; Decoding module 340 is used to decode and generate panoramic images from global features based on learnable query vectors.
[0093] Optionally, the decoding module 340 is specifically used for: dynamically predicting the sampling position and offset of each learnable query vector in the global feature space; sampling and aggregating feature information from the global features based on the sampling position and the offset to generate corresponding panoramic image pixels; wherein each query vector is configured to generate the panoramic image pixels at the corresponding spatial position in the panoramic image.
[0094] Optionally, the aforementioned cross-view feature fusion module 330 is specifically used for: determining a fusion mode for cross-view feature fusion based on system performance requirements; wherein the system performance requirements include high precision requirements and real-time requirements; the fusion mode includes a global fusion mode matching the high precision requirements and a local fusion mode matching the real-time requirements; when the fusion mode is the global fusion mode, establishing an association relationship between the feature representations of all views; when the fusion mode is the local fusion mode, establishing an association relationship between the feature representations of adjacent views; and performing feature fusion based on the established association relationship to obtain the fused global features.
[0095] Optionally, the feature encoding module 320 is specifically used for: determining the azimuth position of each of the images to be stitched in the camera array; wherein, each camera in the camera array is horizontally and equally spaced; generating a relative azimuth position code based on the azimuth position; wherein, the relative azimuth position code is used to characterize the physical spatial topological relationship between each viewpoint; and fusing the relative azimuth position code with the feature encoding result of the corresponding viewpoint to obtain the feature representation of each viewpoint.
[0096] Optionally, the feature encoding module 320 is specifically used to: extract features at different spatial resolution levels to obtain multi-scale feature representations from each viewpoint; wherein, the spatial resolution level includes an original resolution layer and at least one downsampled resolution layer; The aforementioned cross-view feature fusion module 330 is specifically used to: perform cross-view feature fusion based on the correlation between the multi-scale feature representations from different perspectives, and obtain the fused global features.
[0097] Optionally, the image acquisition module 310 is specifically used to: determine a target image from the detection images to be stitched from multiple perspectives in response to a trigger signal; wherein the trigger signal includes an alarm signal generated by a perimeter security system or a panoramic video stitching request signal initiated by a user; and the target image is the detection image to be stitched corresponding to the area indicated by the trigger signal. The aforementioned feature encoding module 320 is specifically used to: encode the features of each target image separately, and embed viewpoint identification information during the encoding process to obtain the feature representation of each viewpoint.
[0098] Optionally, the machine vision-based detection image stitching device 300 further includes: The model training module is used to acquire multi-view training images; synthesize the multi-view training images into corresponding panoramic image labels to construct a training dataset; use the training dataset to train the model and obtain an image stitching model; wherein, the image stitching model is configured to perform a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step.
[0099] Optionally, the machine vision-based detection image stitching device 300 further includes: The video streaming module is used to encode the panoramic image generated by the detection host into a video stream and transmit it to the backend server 130.
[0100] It is understood that the above-mentioned machine vision-based detection image stitching device 300 can realize any one of the functions of the machine vision-based detection image stitching method provided in the embodiments of this application. For the method embodiment section, please refer to the method embodiment section for the way each function is realized and the working principle. The device embodiment section will not repeat it here.
[0101] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of this application. (Refer to...) Figure 4 The electronic device 400 includes a processor 410, a memory 420, and a communication interface 430. These components are interconnected and communicate with each other via a communication bus 440 and / or other forms of connection mechanism (not shown).
[0102] The memory 420 includes one or more (only one is shown in the figure), which may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The processor 410 and other possible components may access the memory 420 to read and / or write data therein.
[0103] Processor 410 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with signal processing capabilities. The processor 410 described above can be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), a network processor (NP), or other conventional processors; it can also be a special-purpose processor, including a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0104] Communication interface 430 includes one or more (only one is shown in the figure) that can be used to communicate directly or indirectly with other devices to exchange data. For example, communication interface 430 can be an Ethernet interface; it can be a mobile communication network interface, such as an interface for 3G, 4G, or 5G networks; or it can be other types of interfaces with data transmission and reception capabilities.
[0105] One or more computer program instructions may be stored in the memory 420. The processor 410 may read and run these computer program instructions to implement the machine vision-based detection image stitching method and other desired functions provided in the embodiments of this application.
[0106] Understandable. Figure 4 The structure shown is for illustrative purposes only; the electronic device 400 may also include more than [other components]. Figure 4 The more or fewer components shown, or having the same Figure 4 The different configurations shown. Figure 4 The components shown can be implemented using hardware, software, or a combination thereof. For example, electronic device 400 can be a single server (or other device with computing power), a combination of multiple servers, a cluster of a large number of servers, etc., and can be either a physical device or a virtual device.
[0107] This application also provides a computer-readable storage medium storing computer program instructions. These instructions are read and executed by a processor to perform the machine vision-based detection image stitching method provided in this application. For example, the computer-readable storage medium can be implemented as follows: Figure 4The memory 420 in the electronic device 400, or a separate storage product (such as a USB flash drive, portable hard drive, etc.).
[0108] This application also provides a computer program product, which includes computer program instructions. These computer program instructions are read and executed by a processor to perform the machine vision-based detection image stitching method provided in this application. For example, these computer program instructions can be stored in... Figure 4 The memory 420 in the electronic device 400 is located inside the memory, or it is stored in a separate storage product (such as a USB flash drive, portable hard drive, etc.).
[0109] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0110] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0111] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0112] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A machine vision-based detection image stitching method, characterized in that, The method includes: Acquire detection images from multiple perspectives to be stitched together; Each of the proposed detection images to be stitched together is feature-encoded, and viewpoint identification information is embedded during the encoding process to obtain the feature representation of each viewpoint; Based on the correlation between the feature representations from different perspectives, cross-perspective feature fusion is performed to obtain the fused global features; A panoramic image is generated by decoding from the global features based on a learnable query vector.
2. The machine vision-based detection image stitching method according to claim 1, characterized in that, The process of decoding and generating a panoramic image from the global features based on a learnable query vector includes: For each learnable query vector, dynamically predict the sampling position and offset of the query vector in the global feature space; Based on the sampling position and the offset, feature information is sampled and aggregated from the global features to generate corresponding panoramic image pixels; wherein, each of the query vectors is configured to generate the panoramic image pixels at the corresponding spatial positions in the panoramic image.
3. The machine vision-based detection image stitching method according to claim 1, characterized in that, The correlation between the feature representations based on different perspectives is used to perform cross-perspective feature fusion to obtain the fused global features, including: Based on system performance requirements, a fusion mode for cross-view feature fusion is determined; wherein, the system performance requirements include high precision requirements and real-time requirements; the fusion mode includes a global fusion mode that matches the high precision requirements, and a local fusion mode that matches the real-time requirements. When the fusion mode is the global fusion mode, establish the association relationship between the feature representations of all views; When the fusion mode is the local fusion mode, the association relationship between the feature representations of adjacent views is established; Feature fusion is performed based on the established relationships to obtain the fused global features.
4. The machine vision-based detection image stitching method according to claim 1, characterized in that, The embedding of viewpoint identification information during the encoding process includes: Determine the azimuth position of each of the images to be stitched in the camera array; wherein, each camera in the camera array is arranged horizontally at equal intervals; A relative azimuth position code is generated based on the azimuth position; wherein, the relative azimuth position code is used to characterize the physical spatial topological relationship between each viewpoint; The relative azimuth position code is fused with the feature code result of the corresponding viewpoint to obtain the feature representation of each viewpoint.
5. The machine vision-based detection image stitching method according to claim 1, characterized in that, The method for obtaining the feature representation includes: extracting features at different spatial resolution levels to obtain multi-scale feature representations from each viewpoint; wherein, the spatial resolution level includes an original resolution layer and at least one downsampled resolution layer; The method of performing cross-perspective feature fusion based on the correlation between the feature representations from different perspectives to obtain fused global features includes: performing cross-perspective feature fusion based on the correlation between the multi-scale feature representations from different perspectives to obtain fused global features.
6. The machine vision-based detection image stitching method according to any one of claims 1-5, characterized in that, The method further includes: In response to a trigger signal, a target image is determined from the detection images to be stitched from multiple perspectives; wherein the trigger signal includes an alarm signal generated by a perimeter security system or a panoramic video stitching request signal initiated by a user; and the target image is the detection image to be stitched corresponding to the area indicated by the trigger signal. The target image is subjected to a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step.
7. The machine vision-based detection image stitching method according to any one of claims 1-5, characterized in that, The method further includes: Acquire training images from multiple perspectives; The multi-view training images are synthesized into corresponding panoramic image labels to construct a training dataset; The model is trained using the training dataset to obtain an image stitching model; wherein the image stitching model is configured to perform a feature encoding step, a cross-view feature fusion step, and a panoramic image generation step.
8. The machine vision-based detection image stitching method according to any one of claims 1-5, characterized in that, The method further includes: The system performs feature encoding, cross-view feature fusion, and panoramic image generation steps locally on the detection host to obtain a panoramic image. The panoramic image is encoded into a video stream and transmitted to the backend server.
9. A machine vision-based detection image stitching device, characterized in that, include: The image acquisition module is used to acquire detection images to be stitched from multiple perspectives; The feature encoding module is used to encode the features of each of the proposed images to be stitched together, and to embed viewpoint identification information during the encoding process to obtain the feature representation of each viewpoint. The cross-view feature fusion module is used to perform cross-view feature fusion based on the correlation between the feature representations from different views, and obtain the fused global features; A decoding module is used to decode and generate a panoramic image from the global features based on a learnable query vector.
10. An electronic device, characterized in that, include: A processor, a memory, and a communication bus, wherein the processor and the memory communicate with each other via the communication bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method as described in any one of claims 1-8 by calling the program instructions.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method as described in any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.