Resolution configuration and spatial position encoding method for multi-camera video and related products
Patent Information
- Application Number
- CN202611248018.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-22
AI Technical Summary
采用统一的固定输出图像尺寸,难以适应不同相机的视频帧尺寸,分别确定各相机的输出图像尺寸,又可能导致实际像素数量之和超过允许的总像素预算,或者使输出图像尺寸与VAE空间下采样因子不匹配
[0035]本申请首先获取当前多相机组合中各相机采集的视频帧、当前多相机组合对应的总像素预算、VAE空间下采样因子以及统一参考网格;其次,根据总像素预算和VAE空间下采样因子,确定当前多相机组合的多相机输出分辨率配置,其中,多相机输出分辨率配置包括当前多相机组合中各相机的输出图像尺寸,各输出图像尺寸的高度和宽度均为VAE空间下采样因子的整数倍,且各输出图像尺寸对应的像素数量之和不超过总像素预算;再次,根据多相机输出分辨率配置,对各视频帧执行空间变换,生成对应的VAE输入图像,并记录空间变换对应的空间变换元数据;然后,对各VAE输入图像进行VAE编码,生成潜在特征,并确定各潜在特征对应的实际潜在空间网格;进一步地,根据统一参考网格、各实际潜在空间网格和空间变换元数据,对各实际潜在空间网格中的潜在空间位置执行参考网格映射,确定各潜在空间位置对应的参考网格连续坐标;最后,根据各潜在空间位置对应的参考网格连续坐标,生成各潜在空间位置对应的空间旋转位置编码,并利用空间旋转位置编码对各潜在特征中相应潜在空间位置处的特征进行位置编码。
Smart Images

Figure CN122802646A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, specifically to a method for resolution configuration and spatial location encoding of multi-camera video and related products. Background Technology
[0002] Robot demonstration data is typically collected by multiple cameras. Due to differences in camera mounting location, sensor specifications, effective field of view, and acquisition parameters, the video frames captured by each camera may have different original image sizes and aspect ratios. Therefore, resolution processing of the video frames from each camera is necessary before inputting them into the video processing model.
[0003] In multi-camera video training and inference, the sum of pixels in the output images from each camera is limited by computational resources. Simultaneously, the output image size needs to be adapted to the VAE spatial downsampling process. Using a uniform, fixed output image size is difficult to adapt to the video frame sizes of different cameras, while determining the output image size for each camera separately may result in the actual sum of pixels exceeding the allowable total pixel budget, or a mismatch between the output image size and the VAE spatial downsampling factor. Therefore, existing processing methods struggle to simultaneously satisfy both the multi-camera total pixel budget constraint and the VAE spatial downsampling constraint.
[0004] Furthermore, different output image sizes, after being encoded by VAE, will form latent spatial grids of different sizes. After video frames are scaled, cropped, or padded, the correspondence between latent spatial locations and original image locations may also change. If spatial location codes are generated directly from integer locations in the latent spatial grid, image locations with the same spatial semantics may correspond to different location coding phases under different latent spatial grids or different spatial transformation conditions, thus causing inconsistencies in location coding scales. Summary of the Invention
[0005] Embodiments of this application propose a resolution configuration and spatial location coding method for multi-camera video, as well as related products.
[0006] In a first aspect, this application provides a method for resolution configuration and spatial location encoding of multi-camera video, the method comprising: Obtain the video frames captured by each camera in the current multi-camera combination, the total pixel budget corresponding to the current multi-camera combination, the VAE spatial downsampling factor, and the unified reference grid; Based on the total pixel budget and the VAE spatial downsampling factor, the multi-camera output resolution configuration of the current multi-camera combination is determined, wherein the multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera combination, the height and width of each output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget; Based on the multi-camera output resolution configuration, spatial transformation is performed on each video frame to generate a corresponding VAE input image, and the spatial transformation metadata corresponding to the spatial transformation is recorded. VAE encoding is performed on each of the VAE input images to generate latent features, and the actual latent space grid corresponding to each latent feature is determined; Based on the unified reference grid, each of the actual potential space grids and the spatial transformation metadata, reference grid mapping is performed on the potential spatial locations in each of the actual potential space grids to determine the continuous reference grid coordinates corresponding to each of the potential spatial locations. Based on the continuous coordinates of the reference grid corresponding to each potential spatial location, a spatial rotation position code is generated for each potential spatial location, and the spatial rotation position code is used to perform position encoding on the features at the corresponding potential spatial locations in each potential feature.
[0007] In some optional implementations, determining the multi-camera output resolution configuration of the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain a pre-established set of finitely valid multi-camera discrete resolution configurations corresponding to the current multi-camera combination and the total pixel budget. The set of finitely valid multi-camera discrete resolution configurations includes multiple sets of candidate multi-camera resolution configurations. Each set of candidate multi-camera resolution configurations includes the candidate output image size of each camera in the current multi-camera combination. The height and width of each candidate output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each candidate output image size does not exceed the total pixel budget. The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor.
[0008] In some optional implementations, determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain the horizontal effective field of view, vertical effective field of view, and preset minimum angle sampling density of each camera in the current multi-camera combination; The basic output image size of each camera is determined based on the horizontal effective field of view, the vertical effective field of view, and the preset minimum angle sampling density. The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the respective base output image sizes, the total pixel budget, and the VAE spatial downsampling factor.
[0009] In some optional implementations, determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain at least one piece of task information from the current action phase, camera roles of each camera, area of the key task region, and visibility of key content; Based on the task information, determine the dynamic budget requirements of each camera in the current multi-camera setup; Based on the dynamic budget requirements and the total pixel budget, the continuous target pixel budget for each camera is determined, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget; The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the continuous target pixel budget and the VAE spatial downsampling factor.
[0010] In some optional implementations, determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Acquire task features and low-cost camera content features corresponding to each camera in the current multi-camera combination; The task features, the content features of each low-cost camera, and the set of finitely legal multi-camera discrete resolution configurations are input into a learnable configuration selection model to generate a configuration selection score for each candidate multi-camera resolution configuration in the set of finitely legal multi-camera discrete resolution configurations. Based on the selected scores for each configuration, the multi-camera output resolution configuration of the current multi-camera combination is determined from the set of finite legal multi-camera discrete resolution configurations.
[0011] In some optional implementations, determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain the original image size of each camera in the current multi-camera setup; Based on the total pixel budget, the continuous target pixel budget for each camera is determined, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget; Calculate the continuous ideal output image size of each camera based on the original aspect ratio corresponding to each original image size and the continuous target pixel budget of each. Based on the continuous ideal output image size, calculate the pixel budget error and aspect ratio error corresponding to each candidate multi-camera resolution configuration in the finite legal multi-camera discrete resolution configuration set; Based on the pixel budget error and the aspect ratio error, the multi-camera output resolution configuration of the current multi-camera combination is jointly determined from the finite set of legal multi-camera discrete resolution configurations.
[0012] In some optional implementations, the step of performing spatial transformation on each of the video frames according to the multi-camera output resolution configuration to generate a corresponding VAE input image, and recording the spatial transformation metadata corresponding to the spatial transformation, includes: Obtain the original image size, full field-of-view preservation requirements, and region of interest information for each video frame; Calculate the corresponding aspect ratio difference based on the original image size and the output image size in the multi-camera output resolution configuration; The reliability of the region of interest is determined based on the region of interest information. Based on the aspect ratio difference, the requirement to preserve the complete field of view, and the reliability of the region of interest, a spatial transformation method is selected, wherein the spatial transformation method includes direct scaling, scaling and then filling, scaling and then safely cropping, or cropping of the region of interest. Perform spatial transformation on each of the video frames according to the spatial transformation method to generate the corresponding VAE input image; Record the spatial transformation parameters corresponding to the spatial transformation method, and generate the spatial transformation metadata.
[0013] In some optional implementations, determining the reliability of the region of interest based on the region of interest information includes: Obtain historical operation region of interest information corresponding to the current action stage and the preceding video frame, and determine the region center, region scale and region confidence corresponding to the current video frame from the operation region of interest information; Based on the current action stage and the region of interest information of the historical operation, time smoothing is performed on the region center and the region scale to determine candidate clipping windows; Verify the coverage of the candidate cropping window with the key content of the task and the preset security boundary; The reliability of the region of interest is determined based on the region confidence level and the coverage.
[0014] In some optional implementations, selecting the spatial transformation method based on the aspect ratio difference, the requirement to preserve the entire field of view, and the reliability of the region of interest includes: Compare the aspect ratio difference with a preset direct scaling threshold; Verify whether the scaled and safe cropping retains the key content of the task; In response to the aspect ratio difference not exceeding the preset direct scaling threshold, direct scaling is selected; In response to the aspect ratio difference exceeding the preset direct scaling threshold and the need to retain the complete field of view, the scaling and filling option is selected. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, and the reliability of the region of interest meeting the preset reliability condition, the region of interest is selected for cropping. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, the reliability of the region of interest not meeting the preset reliability condition, and the safe cropping after scaling retaining the key content of the task, the safe cropping after scaling is selected. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, the reliability of the region of interest not meeting the preset reliability condition, and the scaling-after safe cropping not retaining the key content of the task, the option to revert to scaling-after filling is selected.
[0015] In some optional implementations, the step of performing reference grid mapping on the potential spatial locations in each of the actual potential spatial grids based on the unified reference grid, each of the actual potential spatial grids, and the spatial transformation metadata, to determine the continuous reference grid coordinates corresponding to each of the potential spatial locations, includes: Based on the VAE spatial downsampling factor and each of the actual latent space grids, determine the output image coordinates corresponding to the latent spatial positions in each of the actual latent space grids; Based on the spatial transformation metadata, determine the inverse spatial transformation parameters corresponding to the spatial transformation method, convert each of the output image coordinates into original image coordinates based on the inverse spatial transformation parameters, and determine the valid original image coordinates from the original image coordinates based on the spatial transformation metadata. According to the spatial transformation method, the target mapping coordinates corresponding to each potential spatial location are determined from the output image coordinates and the effective original image coordinates. Based on the target mapping coordinates and the grid size of the unified reference grid, reference grid mapping is performed on each potential spatial location to determine the continuous reference grid coordinates corresponding to each potential spatial location.
[0016] In some optional implementations, determining the target mapping coordinates corresponding to each potential spatial location from the output image coordinates and the effective original image coordinates according to the spatial transformation method includes: In response to the spatial transformation method being direct scaling, the target mapping coordinates corresponding to each potential spatial location are determined based on each actual potential spatial grid and each output image coordinate. In response to the spatial transformation method being any one of the scaling-up fill, scaling-up safe crop, or operation region of interest crop, and without needing to preserve the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial position are determined based on each actual potential spatial grid and each output image coordinate; In response to the spatial transformation method being any one of the scaling-up fill, scaling-up safe crop, or operation region of interest crop, and requiring the preservation of the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial location are determined based on each of the effective original image coordinates and each of the original image sizes.
[0017] In some optional implementations, the step of generating a spatial rotation position code corresponding to each potential spatial position based on the continuous coordinates of the reference grid corresponding to each potential spatial position, and using the spatial rotation position code to perform position encoding on the features at the corresponding potential spatial positions in each potential feature, includes: The rotation position encoding frequency sequence, rotation dimension, and spatial axis channel allocation method used in the strategy model are obtained. Based on the continuous coordinates of the reference grid corresponding to each potential spatial location and the rotation position coding frequency sequence, calculate the height direction rotation phase and width direction rotation phase corresponding to each potential spatial location, and generate the spatial rotation position code corresponding to each potential spatial location; Based on the features at the corresponding potential spatial locations in each of the potential features, generate query vectors and key vectors corresponding to each potential spatial location; Based on the spatial rotation position encoding corresponding to each potential spatial location, the rotation dimension, and the spatial axis channel allocation method, a rotation transformation is performed on the query vector and key vector corresponding to each potential spatial location, and the rotated query vector and key vector are determined as the position encoding result of the feature at the corresponding potential spatial location in each potential feature.
[0018] In some optional implementations, determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain the current task requirements, the current action stage, the action stage of the previous processing moment, and the historical multi-camera output resolution configuration used in the previous processing moment; Based on the total pixel budget, the VAE spatial downsampling factor, and the current task requirements, determine the configuration evaluation value of each candidate multi-camera resolution configuration in the finite legal multi-camera discrete resolution configuration set; Candidate multi-camera output resolution configurations are determined based on the evaluation values of each configuration. In response to the existence of the historical multi-camera output resolution configuration, calculate the configuration gain of the candidate multi-camera output resolution configuration relative to the historical multi-camera output resolution configuration; In response to the absence of the historical multi-camera output resolution configuration, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration, the current action stage being the same as the action stage at the previous processing moment, and the configuration benefit not exceeding a preset switching threshold, the historical multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration and the fact that the current action phase is different from the action phase at the previous processing time, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration, the current action stage being the same as the action stage at the previous processing moment, and the configuration benefit exceeding the preset switching threshold, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination.
[0019] Secondly, this application provides a resolution configuration and spatial location encoding apparatus for multi-camera video, the apparatus comprising: The information acquisition unit is used to acquire video frames captured by each camera in the current multi-camera combination, the total pixel budget corresponding to the current multi-camera combination, the VAE spatial downsampling factor, and the unified reference grid. The resolution configuration determination unit is used to determine the multi-camera output resolution configuration of the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor. The multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera combination. The height and width of each output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget. The spatial transformation unit is used to perform spatial transformation on each of the video frames according to the multi-camera output resolution configuration, generate the corresponding VAE input image, and record the spatial transformation metadata corresponding to the spatial transformation. The VAE encoding unit is used to perform VAE encoding on each of the VAE input images, generate latent features, and determine the actual latent space grid corresponding to each of the latent features; The reference grid mapping unit is used to perform reference grid mapping on the potential spatial locations in each of the actual potential spatial grids based on the unified reference grid, each of the actual potential spatial grids and the spatial transformation metadata, and to determine the continuous reference grid coordinates corresponding to each of the potential spatial locations. The position encoding unit is used to generate a spatial rotation position code corresponding to each potential spatial position based on the continuous coordinates of the reference grid corresponding to each potential spatial position, and to use the spatial rotation position code to perform position encoding on the features at the corresponding potential spatial positions in each potential feature.
[0020] In some optional implementations, the resolution configuration determination unit is further configured to: Obtain a pre-established set of finitely valid multi-camera discrete resolution configurations corresponding to the current multi-camera combination and the total pixel budget. The set of finitely valid multi-camera discrete resolution configurations includes multiple sets of candidate multi-camera resolution configurations. Each set of candidate multi-camera resolution configurations includes the candidate output image size of each camera in the current multi-camera combination. The height and width of each candidate output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each candidate output image size does not exceed the total pixel budget. The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor.
[0021] In some optional implementations, the resolution configuration determination unit is further configured to: Obtain the horizontal effective field of view, vertical effective field of view, and preset minimum angle sampling density of each camera in the current multi-camera combination; The basic output image size of each camera is determined based on the horizontal effective field of view, the vertical effective field of view, and the preset minimum angle sampling density. The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the respective base output image sizes, the total pixel budget, and the VAE spatial downsampling factor.
[0022] In some optional implementations, the resolution configuration determination unit is further configured to: Obtain at least one piece of task information from the current action phase, camera roles of each camera, area of the key task region, and visibility of key content; Based on the task information, determine the dynamic budget requirements of each camera in the current multi-camera setup; Based on the dynamic budget requirements and the total pixel budget, the continuous target pixel budget for each camera is determined, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget; The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the continuous target pixel budget and the VAE spatial downsampling factor.
[0023] In some optional implementations, the resolution configuration determination unit is further configured to: Acquire task features and low-cost camera content features corresponding to each camera in the current multi-camera combination; The task features, the content features of each low-cost camera, and the set of finitely legal multi-camera discrete resolution configurations are input into a learnable configuration selection model to generate a configuration selection score for each candidate multi-camera resolution configuration in the set of finitely legal multi-camera discrete resolution configurations. Based on the selected scores for each configuration, the multi-camera output resolution configuration of the current multi-camera combination is determined from the set of finite legal multi-camera discrete resolution configurations.
[0024] In some optional implementations, the resolution configuration determination unit is further configured to: Obtain the original image size of each camera in the current multi-camera setup; Based on the total pixel budget, the continuous target pixel budget for each camera is determined, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget; Calculate the continuous ideal output image size of each camera based on the original aspect ratio corresponding to each original image size and the continuous target pixel budget of each. Based on the continuous ideal output image size, calculate the pixel budget error and aspect ratio error corresponding to each candidate multi-camera resolution configuration in the finite legal multi-camera discrete resolution configuration set; Based on the pixel budget error and the aspect ratio error, the multi-camera output resolution configuration of the current multi-camera combination is jointly determined from the finite set of legal multi-camera discrete resolution configurations.
[0025] In some optional implementations, the spatial transformation unit is further configured to: Obtain the original image size, full field-of-view preservation requirements, and region of interest information for each video frame; Calculate the corresponding aspect ratio difference based on the original image size and the output image size in the multi-camera output resolution configuration; The reliability of the region of interest is determined based on the region of interest information. Based on the aspect ratio difference, the requirement to preserve the complete field of view, and the reliability of the region of interest, a spatial transformation method is selected, wherein the spatial transformation method includes direct scaling, scaling and then filling, scaling and then safely cropping, or cropping of the region of interest. Perform spatial transformation on each of the video frames according to the spatial transformation method to generate the corresponding VAE input image; Record the spatial transformation parameters corresponding to the spatial transformation method, and generate the spatial transformation metadata.
[0026] In some optional implementations, the spatial transformation unit is further configured to: Obtain historical operation region of interest information corresponding to the current action stage and the preceding video frame, and determine the region center, region scale and region confidence corresponding to the current video frame from the operation region of interest information; Based on the current action stage and the region of interest information of the historical operation, time smoothing is performed on the region center and the region scale to determine candidate clipping windows; Verify the coverage of the candidate cropping window with the key content of the task and the preset security boundary; The reliability of the region of interest is determined based on the region confidence level and the coverage.
[0027] In some optional implementations, the spatial transformation unit is further configured to: Compare the aspect ratio difference with a preset direct scaling threshold; Verify whether the scaled and safe cropping retains the key content of the task; In response to the aspect ratio difference not exceeding the preset direct scaling threshold, direct scaling is selected; In response to the aspect ratio difference exceeding the preset direct scaling threshold and the need to retain the complete field of view, the scaling and filling option is selected. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, and the reliability of the region of interest meeting the preset reliability condition, the region of interest is selected for cropping. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, the reliability of the region of interest not meeting the preset reliability condition, and the safe cropping after scaling retaining the key content of the task, the safe cropping after scaling is selected. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, the reliability of the region of interest not meeting the preset reliability condition, and the scaling-after safe cropping not retaining the key content of the task, the option to revert to scaling-after filling is selected.
[0028] In some optional implementations, the reference mesh mapping unit is further used for: Based on the VAE spatial downsampling factor and each of the actual latent space grids, determine the output image coordinates corresponding to the latent spatial positions in each of the actual latent space grids; Based on the spatial transformation metadata, determine the inverse spatial transformation parameters corresponding to the spatial transformation method, convert each of the output image coordinates into original image coordinates based on the inverse spatial transformation parameters, and determine the valid original image coordinates from the original image coordinates based on the spatial transformation metadata. According to the spatial transformation method, the target mapping coordinates corresponding to each potential spatial location are determined from the output image coordinates and the effective original image coordinates. Based on the target mapping coordinates and the grid size of the unified reference grid, reference grid mapping is performed on each potential spatial location to determine the continuous reference grid coordinates corresponding to each potential spatial location.
[0029] In some optional implementations, the reference mesh mapping unit is further used for: In response to the spatial transformation method being direct scaling, the target mapping coordinates corresponding to each potential spatial location are determined based on each actual potential spatial grid and each output image coordinate. In response to the spatial transformation method being any one of the scaling-up fill, scaling-up safe crop, or operation region of interest crop, and without needing to preserve the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial position are determined based on each actual potential spatial grid and each output image coordinate; In response to the spatial transformation method being any one of the scaling-up fill, scaling-up safe crop, or operation region of interest crop, and requiring the preservation of the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial location are determined based on each of the effective original image coordinates and each of the original image sizes.
[0030] In some optional implementations, the position encoding unit is further configured to: The rotation position encoding frequency sequence, rotation dimension, and spatial axis channel allocation method used in the strategy model are obtained. Based on the continuous coordinates of the reference grid corresponding to each potential spatial location and the rotation position coding frequency sequence, calculate the height direction rotation phase and width direction rotation phase corresponding to each potential spatial location, and generate the spatial rotation position code corresponding to each potential spatial location; Based on the features at the corresponding potential spatial locations in each of the potential features, generate query vectors and key vectors corresponding to each potential spatial location; Based on the spatial rotation position encoding corresponding to each potential spatial location, the rotation dimension, and the spatial axis channel allocation method, a rotation transformation is performed on the query vector and key vector corresponding to each potential spatial location, and the rotated query vector and key vector are determined as the position encoding result of the feature at the corresponding potential spatial location in each potential feature.
[0031] In some optional implementations, the resolution configuration determining unit is further configured to: Obtain the current task requirements, the current action stage, the action stage of the previous processing moment, and the historical multi-camera output resolution configuration used in the previous processing moment; Based on the total pixel budget, the VAE spatial downsampling factor, and the current task requirements, determine the configuration evaluation value of each candidate multi-camera resolution configuration in the finite legal multi-camera discrete resolution configuration set; Candidate multi-camera output resolution configurations are determined based on the evaluation values of each configuration. In response to the existence of the historical multi-camera output resolution configuration, calculate the configuration gain of the candidate multi-camera output resolution configuration relative to the historical multi-camera output resolution configuration; In response to the absence of the historical multi-camera output resolution configuration, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration, the current action stage being the same as the action stage at the previous processing moment, and the configuration benefit not exceeding a preset switching threshold, the historical multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration and the fact that the current action phase is different from the action phase at the previous processing time, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration, the current action stage being the same as the action stage at the previous processing moment, and the configuration benefit exceeding the preset switching threshold, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination.
[0032] Thirdly, this application provides an electronic device, including: one or more processors; a storage device having one or more programs stored thereon; and when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the method described in any embodiment of the first aspect of this application.
[0033] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the first aspect of this application.
[0034] Fifthly, this application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the method described in any embodiment of the first aspect of this application.
[0035] This application first obtains the video frames acquired by each camera in the current multi-camera setup, the total pixel budget corresponding to the current multi-camera setup, the VAE spatial downsampling factor, and the unified reference grid. Second, based on the total pixel budget and the VAE spatial downsampling factor, it determines the multi-camera output resolution configuration of the current multi-camera setup. The multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera setup, where the height and width of each output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget. Third, based on the multi-camera output resolution configuration, it performs a spatial transformation on each video frame to generate the corresponding VAE output. The system first takes an image as input and records the spatial transformation metadata corresponding to the spatial transformation. Then, it performs VAE encoding on each VAE input image to generate latent features and determines the actual latent spatial grid corresponding to each latent feature. Further, based on the unified reference grid, each actual latent spatial grid, and the spatial transformation metadata, it performs reference grid mapping on the latent spatial positions in each actual latent spatial grid to determine the continuous coordinates of the reference grid corresponding to each latent spatial position. Finally, based on the continuous coordinates of the reference grid corresponding to each latent spatial position, it generates the spatial rotation position code corresponding to each latent spatial position and uses the spatial rotation position code to perform position encoding on the features at the corresponding latent spatial positions in each latent feature.
[0036] Through the above technical solution, this application uniformly determines the output image size of each camera in the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor, ensuring that the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget, and that the height and width of each output image size are integer multiples of the VAE spatial downsampling factor. Therefore, it is possible to control the total number of pixels used in multi-camera video processing while adapting the output image size of each camera to the VAE spatial downsampling process, thereby simultaneously satisfying both the multi-camera total pixel budget constraint and the VAE spatial downsampling constraint, and improving the controllability of computing resources.
[0037] Simultaneously, this application records the corresponding spatial transformation metadata when performing spatial transformation on each video frame, and performs reference grid mapping on the potential spatial positions in each actual potential spatial grid by combining a unified reference grid and each actual potential spatial grid. Therefore, it is possible to consider the positional changes caused by different output image sizes and different spatial transformation methods during the reference grid mapping process, and uniformly represent the potential spatial positions in different actual potential spatial grids as continuous coordinates in the same reference coordinate system.
[0038] Based on this, this application generates spatial rotation position codes according to the continuous coordinates of the reference grid corresponding to each potential spatial location, so that locations with the same or similar spatial semantics have comparable position coding phases under different actual potential spatial grids and different spatial transformation conditions. This reduces the impact of changes in the size of the actual potential spatial grid and video frame spatial transformations on the position coding scale, maintains the consistency of the position coding scale under different actual potential spatial grids and different spatial transformation conditions, improves the consistency and stability of the spatial position representation of multi-camera potential features, and is beneficial for subsequent strategy models to uniformly process multi-camera potential features. Attached Figure Description
[0039] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the scope of this application. In the drawings: Figure 1 This is a system architecture diagram of one embodiment of the multi-camera video resolution configuration and spatial location coding method according to this application; Figure 2 This is a flowchart of an embodiment of the multi-camera video resolution configuration and spatial location coding method according to this application; Figure 3 This is an exploded flowchart of an embodiment of step S202 of this application; Figure 4 This is an exploded flowchart of an embodiment of step S203 of this application; Figure 5 This is an exploded flowchart of an embodiment of step S206 of this application; Figure 6 This is a schematic diagram of an embodiment of the multi-camera video resolution configuration and spatial location encoding apparatus according to this application; Figure 7 This is a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of this application. Detailed Implementation
[0040] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0041] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0042] Figure 1 An exemplary system architecture 100 is shown, which can be applied to an embodiment of the multi-camera video resolution configuration and spatial location encoding method of this application.
[0043] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 provides a communication link between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired communication links, wireless communication links, or fiber optic communication links.
[0044] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send data. Terminal devices 101, 102, and 103 can be equipped with multi-camera video processing applications, vision-language-motion model training applications, robot control applications, video data management applications, or other applications related to multi-camera video processing.
[0045] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be desktop computers, portable computers, workstations, training servers, or other electronic devices with data processing capabilities. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. Terminal devices 101, 102, and 103 can be implemented as multiple software programs or software units, or as a single software program or software unit; no specific limitation is made here.
[0046] Server 105 can be a server providing various services, such as a backend server that processes multi-camera video processing requests sent by terminal devices 101, 102, and 103. The backend server can obtain video frames acquired by each camera in the current multi-camera combination, the total pixel budget corresponding to the current multi-camera combination, the VAE spatial downsampling factor, and the unified reference grid. It can then determine the multi-camera output resolution configuration of the current multi-camera combination, perform spatial transformation and VAE encoding on each video frame, and generate spatial rotation position codes based on the unified reference grid. The backend server can also send the latent features after position encoding or the processing results generated based on the latent features after position encoding to terminal devices 101, 102, and 103.
[0047] In some cases, the resolution configuration and spatial location encoding method for multi-camera video provided in this application can be jointly executed by terminal devices 101, 102, and 103 and server 105. For example, the processes of acquiring video frames captured by each camera and performing spatial transformation on each video frame according to the multi-camera output resolution configuration can be performed by terminal devices 101, 102, and 103, while the processes of determining the multi-camera output resolution configuration, performing VAE encoding, performing reference mesh mapping, and generating spatial rotation position encoding can be performed by server 105. This application does not limit this. Correspondingly, the resolution configuration and spatial location encoding device for multi-camera video can also be respectively set in terminal devices 101, 102, and 103 and server 105.
[0048] In some cases, the resolution configuration and spatial location encoding method for multi-camera video provided in this application can be executed by server 105, and correspondingly, the resolution configuration and spatial location encoding device for multi-camera video can also be set in server 105. In this case, system architecture 100 may not include terminal devices 101, 102, and 103.
[0049] In some cases, the resolution configuration and spatial location encoding method for multi-camera video provided in this application can be executed by terminal devices 101, 102, and 103. Correspondingly, the resolution configuration and spatial location encoding device for multi-camera video can also be set in terminal devices 101, 102, and 103. In this case, the system architecture 100 may not include server 105.
[0050] Server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software units, or as a single software program or software unit; no specific limitation is made here.
[0051] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown in the diagram is for illustrative purposes only. Depending on the specific implementation requirements, system architecture 100 may include any number of terminal devices, networks, and servers.
[0052] Continue to refer to Figure 2 , Figure 2 The flowchart S200 of an embodiment of the multi-camera video resolution configuration and spatial location encoding method according to this application is shown.
[0053] Figure 2 The method shown can be derived from Figure 1 The process is executed by the terminal device or server in the process. This process S200 includes the following steps S201 to S206.
[0054] Step S201: Obtain the video frames acquired by each camera in the current multi-camera combination, the total pixel budget corresponding to the current multi-camera combination, the VAE spatial downsampling factor, and the unified reference grid.
[0055] In this embodiment, the current multi-camera combination refers to the set of cameras formed by all enabled cameras participating in multi-camera video processing at the current processing moment. The currently enabled cameras are those in the current multi-camera combination, which may include two or more cameras. Each camera may have a corresponding camera identifier to establish the correspondence between video frames, output image size, latent features, actual latent spatial grid, and spatial transformation metadata during subsequent processing.
[0056] The video frames captured by each camera can be video frames captured by each camera at the same sampling time, or video frames corresponding to the same processing time after time synchronization or time alignment. The terminal device or server can obtain each video frame from each camera, video capture device, local storage device, or video dataset. This application does not limit the specific method of obtaining video frames.
[0057] The total pixel budget for the current multi-camera setup refers to the maximum total number of output pixels allowed to be used by the current multi-camera setup when processing a set of multi-camera video frames corresponding to the current processing time. In other words, the total pixel budget uses a multi-camera sample formed by the current multi-camera setup as the statistical object to constrain the sum of the number of pixels corresponding to the output image size of each camera in the multi-camera sample, and does not represent the time window budget that includes the camera activation frequency.
[0058] The total pixel budget can be determined based on the available computing resources of the electronic device, the video memory capacity, the computational limitations of the VAE, the number of visual features allowed by the policy model, or a pre-defined resource configuration. The total pixel budget can be pre-stored in a system configuration file or dynamically provided by the resource management program based on the currently available computing resources of the electronic device. When the current multi-camera setup changes, the total pixel budget corresponding to the changed multi-camera setup can be obtained.
[0059] The VAE spatial downsampling factor represents the total downsampling factor of the input image in both the height and width directions by the VAE. The VAE spatial downsampling factor can be obtained from the VAE's model structure information, model configuration files, or pre-established model parameter tables. By obtaining the VAE spatial downsampling factor, the output image size of each camera can be adapted to the VAE's spatial downsampling process when subsequently determining the multi-camera output resolution configuration.
[0060] A unified reference grid is used to provide a consistent reference coordinate system for potential spatial locations across different cameras, output image sizes, and actual potential spatial grids. The unified reference grid can include a reference grid height and a reference grid width, and can be pre-stored in a policy model configuration file or a location encoding configuration file.
[0061] In some optional implementations, the unified reference grid can be the standard latent space grid used in the policy model pre-training stage, the latent space grid corresponding to the preset input resolution, or the reference grid corresponding to the preset normalized coordinate interval. Once the unified reference grid is determined, it can be kept consistent during the training and inference processes of the policy model to avoid position encoding phase changes caused by using different reference grids during training and inference.
[0062] Step S201 clarifies the range of cameras currently involved in processing and obtains the basic information required for subsequent determination of multi-camera output resolution configuration, execution of VAE encoding, and reference grid mapping. This provides a data foundation for processing multi-camera video frames under total pixel budget constraints and VAE spatial downsampling constraints, as well as for uniformly representing potential spatial locations in different actual potential spatial grids.
[0063] Step S202: Determine the multi-camera output resolution configuration of the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor.
[0064] The multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera combination. The height and width of each output image size are integer multiples of the VAE space downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget.
[0065] In this embodiment, the multi-camera output resolution configuration is used to uniformly define the output image height and output image width of each camera in the current multi-camera setup. Assume the current multi-camera setup includes... The camera, the first The continuous target pixel budget for each camera is The continuous target pixel budget for each camera satisfies: in, This indicates the number of cameras currently in use in the current multi-camera setup; Indicates the camera number, and The value range is 1 to ; Indicates the first Continuous target pixel budget for each camera; This represents the total pixel budget corresponding to the current multi-camera setup.
[0066] Since the continuous target pixel budget also needs to be converted into discrete output image sizes, the actual number of output pixels after conversion may differ from the continuous target pixel budget. Therefore, the final determined multi-camera output resolution configuration also satisfies: in, and They represent the first The height and width of the output image from each camera; Indicates the first The number of pixels corresponding to the output image size of each camera.
[0067] like Figure 3 As shown, in some optional embodiments, step S202 may include steps S2021 and S2022.
[0068] Step S2021: Obtain a pre-established set of finite legal multi-camera discrete resolution configurations corresponding to the current multi-camera combination and total pixel budget.
[0069] The finite legal multi-camera discrete resolution configuration set includes multiple sets of candidate multi-camera resolution configurations. Each set of candidate multi-camera resolution configurations includes the candidate output image size of each camera in the current multi-camera combination. The height and width of each candidate output image size are integer multiples of the VAE space downsampling factor, and the sum of the number of pixels corresponding to each candidate output image size does not exceed the total pixel budget.
[0070] Let the first Group candidate multi-camera resolution configuration as Then it can be expressed as: in, Indicates the number of the candidate multi-camera resolution configuration, and The value range is 1 to ; This represents the number of candidate multi-camera resolution configurations included in the finite set of legal multi-camera discrete resolution configurations; and They represent the first In the group of candidate multi-camera resolution configurations, the first The height and width of the candidate output images for each camera.
[0071] The candidate output image size in each group of candidate multi-camera resolution configurations satisfies: And satisfy: in, Represents the VAE space downsampling factor; This indicates the modulo operation.
[0072] A finitely valid set of multi-camera discrete resolution configurations can be pre-established based on camera combination identifiers, total pixel budget levels, and VAE spatial downsampling factors, and stored in a configuration file or configuration database. When the current multi-camera combination or total pixel budget changes, the corresponding finitely valid set of multi-camera discrete resolution configurations can be obtained based on the changed camera combination identifiers and total pixel budget.
[0073] Step S2022: Determine the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor.
[0074] Since each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations already satisfies the total pixel budget constraint and the VAE spatial downsampling constraint, a set of candidate multi-camera resolution configurations suitable for the current processing needs can be selected from this set based on at least one of the following: the basic output image size of each camera, task information, low-cost camera content features, or continuous target pixel budget.
[0075] In some optional implementations, the multi-camera output resolution configuration can be determined based on the effective field of view of each camera. Specifically, the horizontal effective field of view, vertical effective field of view, and preset minimum angular sampling density of each camera in the current multi-camera setup are obtained. Based on each horizontal effective field of view, each vertical effective field of view, and the preset minimum angular sampling density, the basic output image size of each camera is determined.
[0076] Let the first The horizontal and vertical effective field of view of each camera are respectively and The preset minimum angle sampling densities in the horizontal and vertical directions are respectively and Then the first The consecutive base output image size of a camera can be expressed as: in, and They represent the first The width and height of the continuous base output image of each camera.
[0077] Here, the continuous base output image size can be upquantized to an integer multiple of the VAE space downsampling factor: in, and They represent the first The base output image width and base output image height of each camera; This represents the function for rounding up.
[0078] After determining the base output image size of each camera, the multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the base output image size, total pixel budget, and VAE spatial downsampling factor.
[0079] Specifically, the total pixel budget and VAE spatial downsampling factor are used as legal constraints for candidate multi-camera resolution configurations, and the dimensions of each basic output image are used as the selection criteria for candidate multi-camera resolution configurations. For each group of candidate multi-camera resolution configurations in the finite set of legal multi-camera discrete resolution configurations, the height difference between the candidate output image height and the corresponding basic output image height of each camera, and the width difference between the candidate output image width and the corresponding basic output image width of each camera are determined. Based on each height difference and each width difference, the basic size difference corresponding to each group of candidate multi-camera resolution configurations is determined. From the candidate multi-camera resolution configurations that satisfy the constraint of an integer multiple of the VAE spatial downsampling factor and whose sum of pixel count does not exceed the total pixel budget, a group of candidate multi-camera resolution configurations whose basic size difference satisfies the preset difference condition is selected, and the selected candidate multi-camera resolution configuration is determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0080] Since each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations already satisfies the integer multiple constraint of the VAE spatial downsampling factor and the total pixel budget constraint, in the actual selection process, the basic size difference between each candidate multi-camera resolution configuration and each basic output image size can be directly compared, and the set of candidate multi-camera resolution configurations with the smallest basic size difference can be selected.
[0081] By adopting the above method, under the premise of satisfying the total pixel budget constraint and VAE space downsampling constraint, the output image size of each camera can be made as close as possible to the basic output image size determined according to the corresponding effective field of view and the preset minimum angle sampling density, so that the multi-camera output resolution configuration takes into account the field of view and minimum angle sampling requirements of each camera.
[0082] In some alternative implementations, the multi-camera output resolution configuration can be determined based on mission information.
[0083] Specifically, acquire at least one piece of task information from the following: current action phase, camera role of each camera, area of key task region, and visibility of key content.
[0084] The current action phase refers to the action execution phase that the robot is in at the current processing moment during the execution of the current task. It is used to characterize the execution purpose and progress of the current action. Depending on the task type, the current action phase may include the search phase, approach phase, alignment phase, grasping phase, manipulation phase, insertion phase, transport phase, or release phase, etc.
[0085] The current action phase can be directly obtained from the robot task control program, the motion state machine, or the task execution state information. For example, when the motion state machine is currently outputting a fine alignment state, the fine alignment phase can be identified as the current action phase. When the robot task control program is currently executing a grasping command, the grasping phase can be identified as the current action phase.
[0086] The camera role is used to indicate the camera's viewpoint or installation location, and can include wrist camera, head camera, top camera, side-view camera, third-person camera, or environmental camera.
[0087] The critical region of a task refers to the image region in the video frames captured by each camera that is directly related to the execution of the current action. It may include at least one of the following: the region where the end effector is located, the region where the manipulated object is located, the region where the target is located, the interaction region between the end effector and the manipulated object, and the environmental context region required to complete the current action.
[0088] The mission-critical region area refers to the image area occupied by the mission-critical region within the corresponding camera's video frame. The mission-critical region area can be represented by the number of pixels included in the mission-critical region, or by the ratio between the mission-critical region area and the total area of the corresponding video frame.
[0089] Key content refers to the image content that needs to be observed to complete the current action, and may include at least one of the following: end effector, manipulated object, target position, and contact relationship between the end effector and the manipulated object. Key content visibility is used to characterize the degree to which key content is observable in the video frame of the corresponding camera.
[0090] Visibility of key content can be represented by discrete visibility states or continuous visibility values. Discrete visibility states can include visible, partially visible, and invisible; continuous visibility values can be determined based on at least one of the following: the area ratio of the key content within the video frame, the occlusion ratio, the target detection confidence, the image sharpness, and whether the key content exceeds the boundaries of the video frame.
[0091] The current action phase, the camera role of each camera, the area of the mission's critical region, and the visibility of critical content can be used individually or in combination to determine the dynamic budget requirements of each camera.
[0092] Based on the mission information, determine the dynamic budget requirements for each camera in the current multi-camera setup.
[0093] Specifically, no. The dynamic budget requirement for a single camera can be expressed as: in, Indicates the first Dynamic budget requirements for each camera; Indicates according to the first The base weights of camera role settings for each camera; This represents the camera importance adjustment factor corresponding to the current action phase; Indicates the critical area of the mission in the first place. Area proportion in video frames of each camera; Indicates the first Visibility of key content corresponding to each camera; This represents a preset function that converts the area of the mission's critical region into budget requirements. This represents a preset function that converts key content visibility into budget requirements.
[0094] After obtaining the dynamic budget requirements for each camera, the continuous target pixel budget for each camera can be determined based on the dynamic budget requirements and the total pixel budget, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget.
[0095] Specifically, it can be expressed as: in, This represents the camera number in the summation. When it is necessary to reserve some pixel budget, it can be... Replace with no more than Allocable pixel budget.
[0096] In some alternative implementations, the dynamic budget requirements for wrist cameras can be increased during fine grasping, alignment, or insertion phases; and the dynamic budget requirements for head cameras, environmental cameras, or third-person cameras can be increased during navigation, search, or long-distance transport phases.
[0097] After determining the budget for each consecutive target pixel, the multi-camera output resolution configuration of the current multi-camera combination can be determined from the finite set of legal multi-camera discrete resolution configurations based on the budget for each consecutive target pixel and the VAE spatial downsampling factor.
[0098] Specifically, for each candidate multi-camera resolution configuration in the set of finite legal multi-camera discrete resolution configurations, the number of candidate output pixels for each camera is determined based on the candidate output image height and candidate output image width of each camera in the candidate multi-camera resolution configuration; the number of candidate output pixels is compared with the continuous target pixel budget of the corresponding camera to determine the pixel budget difference for each camera; the pixel budget differences for each camera in the same candidate multi-camera resolution configuration are accumulated or weighted to determine the budget matching degree corresponding to the candidate multi-camera resolution configuration.
[0099] Select a set of candidate multi-camera resolution configurations from the finite set of legal multi-camera discrete resolution configurations that meet the preset matching conditions for budget matching, and determine the selected candidate multi-camera resolution configurations as the multi-camera output resolution configurations of the current multi-camera combination.
[0100] In some alternative implementations, the set of candidate multi-camera resolution configurations with the smallest pixel budget differences can be determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0101] Since the height and width of each candidate output image in the finite set of legal multi-camera discrete resolution configurations are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each group of candidate multi-camera resolution configurations does not exceed the total pixel budget, the determined multi-camera output resolution configuration can satisfy the VAE spatial downsampling constraint and the total pixel budget constraint, and make the actual number of output pixels of each camera match the corresponding continuous target pixel budget.
[0102] In some alternative implementations, a learnable configuration selection model can also be used to determine the multi-camera output resolution configuration.
[0103] Specifically, it acquires task features and low-cost camera content features corresponding to each camera in the current multi-camera setup.
[0104] Task features refer to features used to characterize at least one of the following information: task objective, task type, action execution progress, and robot current state. Task features provide the learnable configuration selection model with information relevant to the current task requirements, enabling the model to determine the suitability of different candidate multi-camera resolution configurations for the current task. Task features may include at least one of the following: language instruction features, task type features, current action stage features, robot state features, or end effector state features. Task features can be represented using feature vectors, state codes, or structured task parameters.
[0105] Low-cost camera content features refer to the features used to characterize the content of a video frame from a given camera, determined at a lower computational cost than full-resolution VAE encoding, before performing full-resolution VAE encoding on the video frame. Low-cost camera content features can be used to characterize at least one of the following in the corresponding video frame: distribution of task-related objects, visibility of key content, degree of occlusion, degree of motion, or image structure information. Each camera in the current multi-camera setup corresponds to a specific low-cost camera content feature to reflect the differences in the content currently captured by different cameras.
[0106] Low-cost camera content features can be determined based on at least one of the following: low-resolution preview images, historical cached features, lightweight object detection results, image key points, image motion intensity, or occlusion degree. For example, video frames from the corresponding camera can be downsampled, and low-cost camera content features can be extracted using the downsampled video frames; alternatively, camera content features cached at the previous processing time step can be directly read and updated based on the lightweight detection results of the current video frame.
[0107] Task features primarily characterize the overall requirements of the current task, while low-cost camera content features primarily characterize the video frame content currently acquired by each camera. By combining task features and low-cost camera content features, the learnable configuration selection model can simultaneously consider the current task requirements and the differences in video content between different cameras.
[0108] Input the task features, content features of each low-cost camera, and a finite set of legal multi-camera discrete resolution configurations into the learnable configuration selection model to generate configuration selection scores for each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations.
[0109] Specifically, no. The configuration selection score for a group of candidate multi-camera resolution configurations can be expressed as: in, The parameter is Learnable configuration selection model; Indicate task characteristics; This represents the low-cost camera content features corresponding to each camera in the current multi-camera setup; Indicates the first Group candidate multi-camera resolution configuration; Indicates the first Configuration selection score for group of candidate multi-camera resolution configurations.
[0110] After obtaining the configuration selection score of each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations, the multi-camera output resolution configuration of the current multi-camera combination can be determined from the finite set of legal multi-camera discrete resolution configurations based on each configuration selection score.
[0111] Specifically, the configuration selection scores of each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations are compared, the candidate multi-camera resolution configuration with the highest configuration selection score is selected, and the selected candidate multi-camera resolution configuration is determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0112] The learnable configuration selection model directly selects a complete set of candidate multi-camera resolution configurations from a finite set of legal multi-camera discrete resolution configurations, without performing a continuous weighted average of the candidate output image height and candidate output image width in different candidate multi-camera resolution configurations, thereby avoiding the formation of output image sizes that do not meet the VAE spatial downsampling constraints or total pixel budget constraints.
[0113] In some alternative implementations, the multi-camera output resolution configuration can be determined based on the continuous target pixel budget and original aspect ratio of each camera.
[0114] Specifically, obtain the raw image dimensions of each camera in the current multi-camera setup. Determine the consecutive target pixel budget for each camera based on the total pixel budget. The sum of the consecutive target pixel budgets must not exceed the total pixel budget.
[0115] Specifically, a pixel budget allocation ratio can be pre-set for each camera in the current multi-camera setup. The total pixel budget is then allocated to each camera according to its corresponding pixel budget allocation ratio, determining the continuous target pixel budget for each camera. The sum of the pixel budget allocation ratios for each camera does not exceed a preset upper limit, ensuring that the sum of the continuous target pixel budgets does not exceed the total pixel budget.
[0116] In some optional implementations, the pixel budget allocation ratio for each camera can be the same, so that the total pixel budget is evenly distributed among the cameras. In other optional implementations, the pixel budget allocation ratio for each camera can be determined according to a pre-set camera priority, so that higher priority cameras are allocated higher consecutive target pixel budgets. When it is necessary to reserve a portion of the pixel budget, the allocable pixel budget for the current multi-camera combination can be determined first from the total pixel budget, and then the allocable pixel budget can be allocated according to the pixel budget allocation ratio for each camera.
[0117] Calculate the continuous ideal output image size for each camera based on the original aspect ratio corresponding to each original image size and the pixel budget for each continuous target.
[0118] Among them, the The original aspect ratio of a camera can be expressed as: in, and They represent the first The original image width and original image height of each camera; Indicates the first The original aspect ratio of each camera.
[0119] According to the Continuous target pixel budget for each camera and the original aspect ratio The ideal continuous output image size of the camera can be calculated: in, and They represent the first The height and width of the continuous ideal output image of each camera. The continuous ideal output image size satisfies as well as The continuous ideal output image size is used to evaluate candidate multi-camera resolution configurations in a finite set of legal multi-camera discrete resolution configurations, and is not directly used as the input image size for the VAE.
[0120] Based on the size of each continuous ideal output image, calculate the pixel budget error and aspect ratio error corresponding to each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations.
[0121] Among them, the The pixel budget error and aspect ratio error corresponding to the candidate multi-camera resolution configuration can be expressed as: in, Indicates the first The pixel budget error corresponding to the candidate multi-camera resolution configuration; Indicates the first Aspect ratio error corresponding to the resolution configuration of the candidate multi-camera group.
[0122] The multi-camera output resolution configuration of the current multi-camera combination is jointly determined from a finite set of legal multi-camera discrete resolution configurations based on pixel budget error and aspect ratio error.
[0123] Here, based on the pixel budget error and aspect ratio error, non-negative evaluation weights can be assigned to each. The pixel budget error and aspect ratio error corresponding to each candidate multi-camera resolution configuration are then weighted and combined to determine the joint configuration error corresponding to that candidate multi-camera resolution configuration. At least one of the evaluation weights corresponding to the pixel budget error and the aspect ratio error must be greater than zero. These two evaluation weights can be pre-set according to pixel budget matching requirements and aspect ratio preservation requirements.
[0124] The joint configuration error of each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations is compared sequentially, and the candidate multi-camera resolution configuration with the smallest joint configuration error is determined as the multi-camera output resolution configuration of the current multi-camera combination. When the joint configuration error of at least two sets of candidate multi-camera resolution configurations is the same, the candidate multi-camera resolution configuration with the smaller pixel budget error can be selected first; when the pixel budget errors are still the same, one set of candidate multi-camera resolution configurations can be selected according to the preset configuration priority.
[0125] Since each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations satisfies the total pixel budget constraint and the VAE spatial downsampling constraint, the multi-camera output resolution configuration determined by the above method can, while satisfying the above constraints, take into account the degree of matching between the actual number of output pixels and the continuous target pixel budget, as well as the degree of matching between the aspect ratio of the output image and the original aspect ratio.
[0126] In some alternative implementations, the switching of multi-camera output resolution configurations can also be controlled.
[0127] Specifically, the current task requirements, the current action stage, the action stage of the previous processing time, and the historical multi-camera output resolution configuration used in the previous processing time are obtained. Based on the total pixel budget, VAE spatial downsampling factor, and current task requirements, the configuration evaluation value of each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations is determined, and the candidate multi-camera output resolution configuration is determined based on each configuration evaluation value.
[0128] Here, the historical multi-camera output resolution configuration refers to the actual multi-camera output resolution configuration used to process video frames from each camera at the previous processing time, including the output image size adopted by each camera at the previous processing time. The historical multi-camera output resolution configuration is used to characterize the configuration state used before the configuration switch at the current processing time. When there is no previous processing time, or the camera combination at the previous processing time is inconsistent with the current multi-camera combination, causing the original configuration to no longer be applicable, it can be determined that there is no historical multi-camera output resolution configuration.
[0129] The configuration evaluation value is used to characterize the degree of fit between the candidate multi-camera resolution configuration and the current task requirements. The configuration evaluation value can be determined based on at least one of the following: the importance of each camera at the current action stage, the degree to which each camera retains key task content, the degree of matching between the actual number of pixels and the target pixel budget, and the degree of preservation of the output image aspect ratio. In this embodiment, a higher configuration evaluation value indicates a higher degree of fit between the corresponding candidate multi-camera resolution configuration and the current task requirements.
[0130] Based on the total pixel budget, VAE spatial downsampling factor, and current task requirements, the configuration evaluation value of each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations is determined, and the candidate multi-camera resolution configuration with the highest configuration evaluation value is determined as the candidate multi-camera output resolution configuration. When multiple candidate multi-camera resolution configurations have the same configuration evaluation value, the candidate multi-camera output resolution configuration can be determined from them according to the preset configuration priority.
[0131] In response to the existence of historical multi-camera output resolution configurations, calculate the configuration gain of the candidate multi-camera output resolution configuration relative to the historical multi-camera output resolution configuration.
[0132] Among them, configuration benefit is used to characterize the incremental configuration evaluation value that can be generated by switching the historical multi-camera output resolution configuration to the candidate multi-camera output resolution configuration under the current task requirements.
[0133] Specifically, the configuration evaluation value of the historical multi-camera output resolution configuration can be determined based on the current task requirements, and the difference between the configuration evaluation value of the candidate multi-camera output resolution configuration and the configuration evaluation value of the historical multi-camera output resolution configuration can be calculated.
[0134] The configuration gain of the candidate multi-camera output resolution configuration relative to the historical multi-camera output resolution configuration can be expressed as: in, Indicates the return on investment; This represents the configuration evaluation value of the candidate multi-camera output resolution configuration under the current task requirements; This represents the historical multi-camera output resolution configuration evaluation value under the current task requirements.
[0135] When the configuration benefit is positive, it means that the candidate multi-camera output resolution configuration is better suited to the current task requirements than the historical multi-camera output resolution configuration; the greater the configuration benefit, the more significant the improvement in configuration evaluation can be obtained by switching to the candidate multi-camera output resolution configuration.
[0136] The preset switching threshold represents the minimum configuration gain required to allow a multi-camera output resolution configuration switch without any changes in the current action phase. The preset switching threshold can be pre-set based on at least one of the following: preprocessing adjustment overhead caused by resolution configuration switching, VAE input size changes, execution plan regeneration overhead, compiler cache invalidation risk, and configuration stability requirements between adjacent processing times.
[0137] In response to the absence of a historical multi-camera output resolution configuration, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination.
[0138] Since there are no historical configurations available for continued use or comparison at this time, the configuration initialization for the current processing moment can be completed directly using the candidate multi-camera output resolution configuration determined according to the current task requirements.
[0139] In response to the existence of a historical multi-camera output resolution configuration, the current action stage being the same as the action stage at the previous processing moment, and the configuration gain not exceeding a preset switching threshold, the historical multi-camera output resolution configuration is determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0140] The fact that the current action phase remains unchanged indicates that the task focus at adjacent processing times typically does not change substantially. The fact that the configuration gain does not exceed the preset switching threshold indicates that the improvement in the suitability of the candidate multi-camera output resolution configuration relative to the historical multi-camera output resolution configuration is insufficient to trigger a configuration switch. Therefore, continuing to use the historical multi-camera output resolution configuration can reduce unnecessary configuration switches that occur when the gain is small.
[0141] In response to the existence of historical multi-camera output resolution configurations and the fact that the current action phase is different from the action phase at the previous processing time, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination.
[0142] A change in the current action phase indicates that the processing focus of the current task or the importance of different cameras may have changed, and the historical multi-camera output resolution configuration may no longer be suitable for the changed task requirements. Therefore, the candidate multi-camera output resolution configuration determined according to the current task requirements can be directly adopted without being restricted by the preset switching threshold.
[0143] In response to the existence of historical multi-camera output resolution configurations, the current action phase being the same as the action phase at the previous processing moment, and the configuration benefit exceeding a preset switching threshold, the candidate multi-camera output resolution configuration is determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0144] Although the current action phase remains unchanged, the configuration benefit exceeds the preset switching threshold, indicating that the candidate multi-camera output resolution configuration can provide a sufficient improvement in adaptability to trigger a configuration switch. Therefore, using the candidate multi-camera output resolution configuration at the current processing moment allows the output image size of each camera to better adapt to the current task requirements.
[0145] When changes occur in the current multi-camera setup, rendering the historical multi-camera output resolution configuration no longer applicable, the historical multi-camera output resolution configuration can be considered non-existent. Changes in the multi-camera setup can include the addition, removal, or deactivation of cameras, or changes in camera identifiers. Since the cameras and their output image sizes included in the historical multi-camera output resolution configuration cannot be matched one-to-one with the changed current multi-camera setup, this historical configuration is no longer used as the benchmark for configuration switching. Instead, the current configuration is determined as if no historical multi-camera output resolution configuration exists.
[0146] By controlling the switching of multi-camera output resolution configuration based on changes in action phases and configuration benefits, the output image size can be updated in a timely manner when the focus of task processing changes or when candidate configurations can produce significant configuration benefits. When the focus of task processing does not change and the configuration benefits are small, the historical configuration can continue to be used. This reduces unnecessary changes in output image size between adjacent processing moments, improves the stability of VAE input size, and facilitates the reuse of existing execution plans or compilation caches.
[0147] It should be noted that the configuration determination methods based on effective field of view, task information, learnable configuration selection models, and continuous target pixel budget and original aspect ratio can be used individually or in combination according to actual processing needs. When using them in combination, the continuous target pixel budget of each camera can be determined first using the effective field of view or task information, and then the complete configuration can be selected from a finite set of legal multi-camera discrete resolution configurations using pixel budget error and aspect ratio error; alternatively, the rule-determined configuration can be used as a fallback configuration when the learnable configuration selection model fails.
[0148] Step S202 allows for the determination of a complete multi-camera output resolution configuration that satisfies the VAE spatial downsampling constraint within the total pixel budget of the current multi-camera combination. This avoids the situation where the sum of the actual pixel count exceeds the total pixel budget after determining the output image size of each camera separately. Simultaneously, finite discrete configurations reduce the number of arbitrary input sizes encountered during training and inference. Configuration selection based on effective field of view, task information, or camera content features ensures that the output image size is adapted to the field of view of different cameras and the current processing requirements. Configuration switching control improves the stability of input sizes between adjacent processing moments.
[0149] Step S203: Based on the multi-camera output resolution configuration, perform spatial transformation on each video frame to generate the corresponding VAE input image, and record the spatial transformation metadata corresponding to the spatial transformation.
[0150] In this embodiment, spatial transformation is used to convert video frames with different original image sizes and different original aspect ratios into output image sizes specified by the multi-camera output resolution configuration.
[0151] For each camera in the current multi-camera setup, a corresponding spatial transformation method can be determined, and the video frames acquired by that camera can be processed according to the determined spatial transformation method. The height and width of the generated VAE input image are equal to the output image height and output image width determined for the corresponding camera in the multi-camera output resolution configuration, respectively.
[0152] Spatial transformation metadata is used to record the spatial transformation method and parameters used to transform video frames from the original image space to the VAE input image space, so as to determine the correspondence between the VAE input image position and the original image position in subsequent processing.
[0153] refer to Figure 4 , Figure 4 This is an exploded flowchart of an embodiment of step S203 of this application.
[0154] like Figure 4 As shown, in some optional embodiments, step S203 may include steps S2031 to S2036.
[0155] Step S2031: Obtain the original image size, full field of view preservation requirements, and region of interest information for each video frame.
[0156] The original image size includes the original image height and original image width of the video frame. The complete field-of-view preservation requirement indicates whether the entire effective field of view contained in the original video frame should be preserved in the VAE input image after spatial transformation. The complete field-of-view preservation requirement can be determined based on at least one of the following: current task type, current action stage, camera role, and whether the edge regions of the original image contain task-critical content.
[0157] For example, during the navigation, search, or environmental awareness phases, the edge regions of the original image may contain obstacles, target objects, or environmental structures, thus determining that the entire field of view needs to be preserved; during the fine grasping, alignment, or interpolation phases, when the key content of the task is concentrated in a local operation area, it can be determined that the entire field of view does not need to be preserved.
[0158] Operational region of interest (ROI) information is used to characterize the image region directly related to the current robot operation. It may include at least one of the following: end effector, manipulated object, operation target location, interaction area, and environmental context required to complete the current operation. ROI information can be determined through object detection results, image segmentation results, keypoint detection results, robot pose projection results, language command localization results, historical spatial attention results, or a combination of the above.
[0159] The region of interest (ROI) information may include at least one of the following: region center, region scale, region boundary, region confidence score, and region generation source. Specifically, the region center represents the central location of the ROI within the original video frame; the region scale represents the coverage area of the ROI; and the region confidence score represents the reliability of the ROI localization result.
[0160] Step S2032: Calculate the corresponding aspect ratio difference based on the original image size and the output image size in the multi-camera output resolution configuration.
[0161] Let the first The aspect ratios of the original images, the aspect ratios of the output images, and the aspect ratio differences of the cameras are as follows: , and Then it can be expressed as: in, and They represent the first The original image width and original image height of each camera; and They represent the first The output image width and output image height of each camera; Indicates the first The aspect ratio differences between the cameras.
[0162] The smaller the aspect ratio difference, the smaller the non-proportional deformation caused when the original video frame is directly scaled to the corresponding output image size; the larger the aspect ratio difference, the more obvious the changes that direct scaling may cause to the shape and spatial structure of objects in the video frame.
[0163] Step S2033: Determine the reliability of the region of interest based on the region of interest information.
[0164] The reliability of the region of interest (ROI) indicates whether the ROI in the current video frame can reliably serve as the basis for determining the cropping window. If the ROI reliability meets the preset reliability conditions, cropping can be performed based on the ROI; if the ROI reliability does not meet the preset reliability conditions, cropping will not be performed based on the current ROI, thereby reducing the risk of cropping critical content due to region positioning errors, region jitter, or insufficient region coverage.
[0165] In some alternative implementations, determining the reliability of the region of interest based on the region of interest information may include the following processing.
[0166] Obtain historical operation region of interest information corresponding to the current action stage and previous video frames, and determine the region center, region scale, and region confidence of the current video frame from the operation region of interest information.
[0167] Among them, the preceding video frame refers to the video frame captured by the current camera before the current video frame; where, the first The video frame represents the first video frame. A video frame is the preceding video frame that is temporally adjacent to the current video frame. The preceding video frame includes one or more video frames captured by the current camera before the current video frame.
[0168] Historical region of interest (ROI) information refers to the ROI information determined in one or more video frames preceding the current video frame. It may include at least one of the following: historical region center, historical region scale, historical region confidence level, and historical cropping window. Historical ROI information is used to determine whether the changes in the current ROI are continuous over time and to reduce positional and scale jitter in cropping windows between adjacent video frames.
[0169] Let the first The region of interest for each video frame is Then it can be expressed as: in, and They represent the first The x and y coordinates of the center of the region of interest for each video frame in the original image coordinate system; Indicates the first The region scale corresponds to a video frame. The region scale can be represented by the crop window height, crop window width, region area, or region diagonal length.
[0170] In one alternative approach, the cropping window height can be used to represent the region scale, and the cropping window width can be determined based on the aspect ratio of the output image, ensuring that the aspect ratio of the candidate cropping window matches that of the output image. The height and width of the candidate cropping window can be expressed as: in, and They represent the first The candidate cropping window height and width are set for each video frame. Ensuring the aspect ratio of the candidate cropping window matches the aspect ratio of the output image reduces non-proportional distortion that occurs when scaling the region of interest again after cropping.
[0171] Based on the current action stage and historical operation information about the region of interest, time smoothing is performed on the region center and the region scale to determine candidate clipping windows.
[0172] Temporal smoothing integrates the region of interest (ROI) information of the current video frame with the historical ROI information of previous video frames. When the current action phase remains unchanged, it enhances the role of historical ROI information in temporal smoothing to maintain a stable cropping window. When the current action phase changes, it enhances the role of the ROI information of the current video frame in temporal smoothing, enabling the cropping window to promptly follow changes in the critical regions of the task.
[0173] The region center can be processed using exponential smoothing, and the region scale can be smoothed in logarithmic space. The smoothed region center and region scale can be expressed as: in, and Indicates the center of the smoothed region; Represents the logarithmic scale after smoothing; Indicates the scale of the smoothed region; The current frame following weight indicates the center of the region; The current frame following weight represents the region scale; , and This indicates the smoothing result determined based on the preceding video frames.
[0174] The current frame following weight at the region center and the current frame following weight at the region scale can satisfy... This ensures that changes in the region scale do not outpace changes in the region center, thereby reducing significant scaling of the cropping window between adjacent video frames. When historical region of interest information is unavailable, the smoothing result can be initialized using the region center and scale corresponding to the current video frame.
[0175] Candidate cropping windows can be determined based on the smoothed region center and the smoothed region scale. The candidate cropping window should be located within the original image area; when the candidate cropping window exceeds the original image boundary, it can be moved, reduced in size, or the excess portion can be compensated for.
[0176] Verify the coverage of the candidate clipping window with the key content of the task and the preset safety boundaries.
[0177] The key elements of the mission can include the end effector, the manipulated object, the target location, and the environmental context required to complete the current action. The preset safety boundary refers to the extended range added outside the area corresponding to the key elements of the mission, used to reserve space for target motion, detection errors, time smoothing lag, and camera shake.
[0178] Let the region of interest for operation after adding a preset safety boundary be... ⊕m, the candidate cropping window is Then the coverage condition can be expressed as: Where m represents the preset safety boundary; ⊕m represents expanding the region of interest outward from the preset safety boundary; ⊆ represents the expanded region of interest being covered by the candidate clipping window.
[0179] When a candidate cropping window cannot cover the critical content of the task and the preset safety boundary, the candidate cropping window can be moved or expanded until the coverage requirement is met; when a candidate cropping window that meets the coverage requirement cannot be formed within the original image range, it can be determined that the reliability of the region of interest does not meet the preset reliability condition.
[0180] The reliability of the region of interest is determined based on the regional confidence level and coverage.
[0181] Specifically, the region confidence level can be compared with a preset confidence threshold to determine whether the candidate pruning window meets the coverage condition. When the region confidence level is not lower than the preset confidence threshold, and the candidate pruning window can cover the key content of the task and the preset security boundary, it can be determined that the reliability of the region of interest meets the preset reliability condition; when the region confidence level is lower than the preset confidence threshold, or the candidate pruning window cannot cover the key content of the task and the preset security boundary, it can be determined that the reliability of the region of interest does not meet the preset reliability condition.
[0182] In some optional methods, the temporal consistency of the region of interest (ROI) can be verified based on the degree of overlap between the current ROI and the current region predicted using historical ROIs. Only when the region confidence, coverage, and temporal consistency all meet their respective conditions is the reliability of the ROI determined to meet the preset reliability criteria. Saliency maps or attention maps can be used to assist in generating ROIs, but they are not used as the sole basis for performing ROI pruning if the region confidence and coverage checks are not passed.
[0183] Step S2034: Select a spatial transformation method based on the aspect ratio difference, the requirement to preserve the complete field of view, and the reliability of the region of interest.
[0184] The spatial transformation methods include direct scaling, scaling followed by filling, scaling followed by safe cropping, or cropping the region of interest.
[0185] Direct scaling refers to scaling the original video frame to the corresponding output image size along both the image height and width directions. Direct scaling can completely preserve the image content of the original video frame, but when the aspect ratios of the original image and the output image are different, non-proportional distortion may occur.
[0186] Post-scale padding refers to scaling the original video frame proportionally according to the original image's aspect ratio, so that the scaled, complete image is within the output image's range, and a preset pixel value is filled into the remaining area between the scaled image and the output image's boundary. Post-scale padding preserves the complete field of view of the original video frame, but it will create a filled area in the VAE input image that does not correspond to the content of the original image.
[0187] Scale-safe cropping refers to scaling the original video frames proportionally to their original aspect ratio, ensuring the scaled image covers the target canvas corresponding to the output image, and then cropping out any edges that extend beyond the target canvas according to a preset safe cropping window. The safe cropping window can be determined based on the camera's role, camera mounting position, or preset geometric rules, and must ensure that critical content is preserved after cropping.
[0188] Operational region of interest (ROI) cropping refers to determining a cropping window based on a region of interest that meets preset reliability conditions, extracting a local image corresponding to the cropping window from the original video frame, and scaling the extracted local image to the corresponding output image size. ROI cropping allows a limited number of output pixels to be concentrated on representing image content directly related to the current operation.
[0189] When selecting a spatial transformation method, compare the aspect ratio difference with the preset direct scaling threshold.
[0190] The preset direct scaling threshold is used to limit the extent of aspect ratio change allowed by direct scaling. It can be predetermined based on the tolerance of image deformation by the pre-set strategy model, the task type, or the verification data.
[0191] For example, the preset direct scaling threshold can be set to any value between 0.01 and 0.03, but this range does not constitute a limitation on this application.
[0192] At the same time, verify whether the safe cropping after scaling preserves the key content of the task.
[0193] Specifically, based on the scaling ratio and a preset safe cropping window, the range of the original image to be retained after scaling and safe cropping can be determined, and it can be determined whether the end effector, the manipulated object, the target location, and the necessary environmental context are within this original image range. If all the critical content required by the task is within this original image range, it is determined that the critical content of the task will be retained after scaling and safe cropping; otherwise, it is determined that the critical content of the task will not be retained after scaling and safe cropping.
[0194] If the aspect ratio difference does not exceed the preset direct scaling threshold, direct scaling is selected.
[0195] At this point, the difference between the aspect ratio of the original image and the aspect ratio of the output image is small, and the non-proportional deformation caused by direct scaling is within the allowable range. Therefore, the spatial transformation process can be simplified while preserving the complete image content.
[0196] In response to aspect ratio differences exceeding the preset direct scaling threshold and the need to preserve the full field of view, select to scale and then fill.
[0197] At this point, direct scaling may produce image distortions that exceed the allowable range, and the edge content of the original image cannot be discarded by cropping. Therefore, scaling and padding are used to preserve the complete effective field of view of the original video frame.
[0198] In response to situations where the aspect ratio difference exceeds a preset direct scaling threshold, there is no need to retain the entire field of view, and the reliability of the region of interest meets the preset reliability conditions, the region of interest is selected for cropping.
[0199] At this point, the range of the original image that needs to be retained can be narrowed based on the reliable region of interest for the operation, so that more pixels in the output image can be used to represent the key content required for the current operation.
[0200] In response to aspect ratio differences, when the aspect ratio exceeds the preset direct scaling threshold, when it is not necessary to retain the complete field of view, when the reliability of the region of interest does not meet the preset reliability conditions, and when safe cropping after scaling retains the key content of the task, select safe cropping after scaling.
[0201] At this point, although the dynamic cropping window cannot be reliably determined based on the current image content, a preset safe cropping window can be used to crop out some edge areas while ensuring that the mission-critical content is not cropped.
[0202] In response to situations where the aspect ratio difference exceeds the preset direct scaling threshold, there is no need to retain the complete field of view, the reliability of the area of interest does not meet the preset reliability conditions, and the scaling and safe cropping does not retain the task's critical content, the option to roll back and fill after scaling is selected.
[0203] At this point, both cropping the region of interest and scaling and then safely cropping could result in the loss of critical content. Therefore, scaling and then padding preserves the complete view of the original video frame.
[0204] Step S2035: Perform spatial transformation on each video frame according to the spatial transformation method to generate the corresponding VAE input image.
[0205] When the spatial transformation method is direct scaling, the horizontal scaling ratio can be determined by the ratio between the output image width and the original image width, and the vertical scaling ratio can be determined by the ratio between the output image height and the original image height. Then, interpolation processing is performed on the original video frames based on the horizontal and vertical scaling ratios. The horizontal and vertical scaling ratios can be different.
[0206] When the spatial transformation method is scaled and padded, the scaling ratio can be determined according to the principle of ensuring that the complete original video frame enters the output image range. Specifically, the smaller of the ratio of the output image width to the original image width and the ratio of the output image height to the original image height can be selected as the scaling ratio. After scaling the original video frame proportionally, padded areas can be added to the top, bottom, left, or right of the scaled image to make the processed image reach the corresponding output image size. The padded pixels can be determined using a preset constant value, the image mean, or the edge pixel expansion value.
[0207] When the spatial transformation method is scaled and safely cropped, the scaling ratio can be determined according to the principle of ensuring that the scaled image covers the output image range. Specifically, the larger of the ratio of the output image width to the original image width and the ratio of the output image height to the original image height can be selected as the scaling ratio. After scaling the original video frames proportionally, the edge portions that exceed the output image range are cropped according to the preset safe cropping window to generate the corresponding VAE input image.
[0208] When the spatial transformation method is region-of-interest (ROI) cropping, a local image can be extracted from the original video frame based on the candidate cropping window, and then the local image can be scaled to the corresponding output image size. If the candidate cropping window exceeds the boundaries of the original image, its position or scale can be adjusted to ensure it remains within the original image's range. For consecutive video frames, time-smoothed candidate cropping windows can be used to perform RIO cropping, reducing spatial jitter between adjacent VAE input images.
[0209] Step S2036: Record the spatial transformation parameters corresponding to the spatial transformation method and generate spatial transformation metadata.
[0210] Spatial transformation metadata can be recorded separately for each camera and each video frame, and may include at least one of the following: camera identifier, video frame identifier, original image size, output image size, spatial transformation method, scaling ratio, crop window, padding range, and effective region mask.
[0211] When the spatial transformation method is direct scaling, the spatial transformation metadata can at least record the horizontal scaling ratio and the vertical scaling ratio to characterize the non-uniform scaling relationship between the original image coordinates and the output image coordinates.
[0212] When the spatial transformation method is scaling followed by padding, the spatial transformation metadata can at least record the proportional scaling ratio, the scaled image size, the padding width in each direction, and the effective region mask. The effective region mask is used to distinguish between the effective regions formed by the original image content and the invalid regions formed by the padding pixels in the output image.
[0213] When the spatial transformation method is scaled-safe cropping, the spatial transformation metadata can at least record the scaling ratio, the size of the scaled image, and the position and size of the cropping window in the scaled image, so as to restore the corresponding position of the output image in the original image based on the cropping offset.
[0214] When the spatial transformation method is to crop the region of interest, the spatial transformation metadata can at least record the position and size of the cropping window in the original image, as well as the scaling ratio used to scale the cropped image to the output image size in the horizontal and vertical directions.
[0215] In some optional methods, the spatial transformation metadata can also record the coordinate origin, pixel center definition, interpolation method, boundary processing method, and configuration version used for pixel coordinates, so as to ensure that subsequent reference mesh mapping adopts the same coordinate convention as the spatial transformation.
[0216] Step S203 allows for the selection of a suitable spatial transformation method among direct scaling, post-scaling padding, post-scaling safe cropping, and region-of-interest (ROI) cropping, based on the aspect ratio difference between the original video frame and the target output image, the requirement to preserve the entire field of view, and the reliability of the ROI. This ensures that the entire camera field of view is preserved when needed, while also allowing output pixels to be concentrated on mission-critical content when the ROI is reliable. Simultaneously, by recording the scaling ratio, cropping window, padding range, and effective region mask, the correspondence between the VAE input image position and the original image position can be preserved, providing a spatial transformation basis for subsequent reference grid mapping and spatial rotation position encoding.
[0217] Step S204: Perform VAE encoding on each VAE input image to generate latent features and determine the actual latent space grid corresponding to each latent feature.
[0218] In this embodiment, VAE encoding refers to using the encoder in a variational autoencoder to extract features and spatially downsample the VAE input image, converting the image in pixel space into a feature representation in latent space. Latent features refer to the feature tensor output by the VAE encoder in the latent space, which may include channel dimension, latent space height dimension, and latent space width dimension.
[0219] The actual latent space grid refers to a two-dimensional location grid established based on the actual latent space height and width of the latent features. This grid is used to describe the arrangement of spatial positions in the latent features, and its grid size depends on the actual latent feature size output by the VAE encoder for the current VAE input image, rather than a pre-assumed fixed size.
[0220] Specifically, the VAE input images corresponding to each camera in the current multi-camera combination are input into the VAE encoder, so that the VAE encoder can perform spatial downsampling and feature extraction on each VAE input image.
[0221] Let the first The camera in the The VAE input image corresponding to each video frame is The VAE encoder is Then the first The camera in the The latent features corresponding to each video frame can be represented as: in, Indicates the camera number; Indicates the time number of the video frame; Indicates the first The camera in the VAE input image corresponding to each video frame; represents the parameter VAE encoder; represents the -th camera in the -th video frame corresponding latent feature.
[0222] The VAE encoder can output parameters of a latent variable distribution and obtain latent features according to the latent variable distribution; it may also use the mean of the latent variable distribution or a result obtained by performing a preset scaling process on the mean as the latent feature during inference. The present application does not limit the specific generation manner of latent features, as long as the VAE encoder can convert a VAE input image into a latent feature with a determined spatial size.
[0223] VAE input images corresponding to respective cameras in the current multi-camera combination can be respectively encoded by the same VAE encoder. Encoding performed by the same VAE encoder can enable latent features corresponding to respective cameras to have consistent channel meanings and spatial downsampling rules. Latent features corresponding to different cameras can have different actual latent space heights or actual latent space widths, but their channel dimensions and the feature meanings represented by respective channels remain consistent.
[0224] After obtaining each latent feature, read size information of each latent feature tensor to obtain an actual spatial size of each latent feature. Taking the latent feature corresponding to the -th camera as an example, the actual spatial size of the latent feature comprises an actual latent space height and an actual latent space width. The actual latent space height and the actual latent space width are subject to the latent feature tensor size actually output by the VAE encoder, thereby avoiding inconsistency between the number of positions and the latent feature caused by using a preset fixed latent space size.
[0225] When the VAE encoder adopts the same VAE spatial downsampling factor in the height direction and the width direction , and does not adopt special padding or boundary processing that changes a conventional size mapping relationship, the actual latent space height and the actual latent space width can be expressed as: wherein, represents the actual latent space height corresponding to the -th camera; represents the actual latent space width corresponding to the -th camera; and respectively represent the VAE input image height and the VAE input image width corresponding to the -th camera; represents the VAE spatial downsampling factor.
[0226] In step S202, the height and width of each output image size are determined to be integer multiples of the VAE spatial downsampling factor. Therefore, when using the above-mentioned conventional size mapping relationship, the actual latent space height and actual latent space width are both integers. This reduces the additional padding, truncation, or size rounding introduced by the mismatch between the VAE input image size and the VAE spatial downsampling structure.
[0227] When the VAE encoder uses different spatial downsampling factors in the height and width directions, the actual latent space height and width can be determined based on the spatial downsampling factors in the height and width directions, respectively. When the VAE encoder uses asymmetric padding, special boundary processing, or other size mapping methods, the latent space size can be obtained according to the size mapping rules corresponding to the VAE encoder, and the final size is based on the latent feature tensor size actually output by the VAE encoder.
[0228] Based on the actual potential space height and actual potential space width of each potential feature, establish the actual potential space grid corresponding to each potential feature. The actual potential space grid corresponding to each camera can be represented as: in, Indicates the first The actual potential space grid corresponding to each camera; This represents the height-direction position index in the actual potential space grid; This represents the position index in the width direction within the actual potential space grid; and They represent the first The actual potential space height and actual potential space width corresponding to each camera.
[0229] Each grid location in the actual latent spatial grid corresponds to a spatial feature vector in the latent features. For example, grid location Corresponding to the Potential features of each camera in the height orientation position index and width direction position index The feature vector at a given location. The number of grid locations included in the actual latent space grid is equal to the product of the actual latent space height and the actual latent space width. Therefore, there is a one-to-one correspondence between the actual latent space grid and the spatial location of the corresponding latent feature.
[0230] It should be noted that the actual latent space grid is determined based on the actual feature size obtained after VAE encoding of the current VAE input image, and its grid size can vary with the output image size of different cameras. The actual latent space grid is used to describe the true spatial arrangement of the corresponding latent features; the unified reference grid is used to provide a uniform coordinate scale for positions in different actual latent space grids, and the two serve different purposes. Subsequently, the positions in each actual latent space grid can be mapped to the unified reference grid without rescaling latent features of different spatial sizes to the same size, thereby reducing the loss of feature information caused by secondary interpolation of latent features.
[0231] Step S204 allows for the acquisition of corresponding latent features from the VAE input images for each camera. An actual latent space grid is then established based on the actual spatial dimensions of these features, ensuring that the number and arrangement of positions within the grid match the actual spatial locations of the latent features. Using the actual spatial dimensions output by the VAE encoder as a basis also allows for compatibility with different VAE size mapping rules and provides an accurate position index foundation for subsequently mapping the latent spatial locations in each actual latent space grid to a unified reference grid.
[0232] Step S205: Based on the unified reference grid, each actual potential space grid and spatial transformation metadata, perform reference grid mapping on the potential spatial locations in each actual potential space grid to determine the continuous coordinates of the reference grid corresponding to each potential spatial location.
[0233] The actual potential spatial grid refers to the two-dimensional location grid determined in step S204 based on the actual output feature dimensions of the VAE. Potential spatial location refers to the height and width positions within this two-dimensional location grid, which can be indexed by the height position. and width direction position index The unified reference grid refers to a standard two-dimensional coordinate scale commonly used by different cameras, output resolutions, and spatial transformation methods. Its height and width are denoted as […]. and .
[0234] Continuous coordinates of the reference grid refer to the coordinates obtained by mapping the latent spatial location to a unified reference grid. These coordinates can be integers or non-integers, and are not limited to discrete grid indices within the unified reference grid. The reference grid mapping is only used to unify the coordinate representation of the latent spatial location and does not perform scaling, interpolation, or resampling on the latent features themselves.
[0235] The unified reference grid can be determined based on the standard latent space grid used in the policy model pre-training phase, the latent space grid corresponding to the preset standard output image size, or a preset coordinate range. The grid size and coordinate conventions of the unified reference grid can be saved as model configuration and maintained consistently during the training and inference phases.
[0236] In some optional implementations, step S205 may sequentially determine the output image coordinates, determine the valid original image coordinates using inverse spatial transformation parameters, and select the target mapping coordinates according to the spatial transformation method and perform reference grid mapping.
[0237] First, based on the VAE spatial downsampling factor and each actual latent space grid, determine the output image coordinates corresponding to the latent spatial locations in each actual latent space grid.
[0238] The output image coordinates here refer to the representative coordinates of the potential spatial location in the VAE input image obtained in step S203.
[0239] Let the first The height and width of the actual potential space grid corresponding to each camera are respectively... and The potential spatial location is represented as Using a pixel index coordinate system, the center coordinates of the first pixel at the top left corner of the output image are set to... Furthermore, when the VAE encoder uses regular space downsampling, the output image coordinates corresponding to this latent spatial location can be expressed as: in, (w) indicates the first The x-coordinate of the output image corresponding to the width direction position index w in the actual latent space grid of each camera; (h) represents the vertical coordinate of the output image corresponding to the height position index h; This represents the VAE space downsampling factor.
[0240] The above formula uses the center of the output image region corresponding to each potential spatial location as the representative location, where half is subtracted to convert the continuous coordinates with the image boundary as the origin into pixel index coordinates with the center of the first pixel in the upper left corner as the origin. Therefore, a potential spatial location corresponds to the center of its covered area in the output image, rather than to the boundary of the covered area.
[0241] When the spatial step size, effective receptive field center, or boundary processing method of a specific VAE encoder differs from the above formula, the output image coordinates can be determined based on the actual structure of the VAE encoder, and the corresponding coordinate conversion parameters can be saved as VAE configuration. The position index in the actual latent space grid satisfies as well as .
[0242] Secondly, based on the spatial transformation metadata, the inverse spatial transformation parameters corresponding to the spatial transformation method are determined, and the coordinates of each output image are converted into the coordinates of the original image based on the inverse spatial transformation parameters.
[0243] Spatial transformation metadata may include at least one of the following: spatial transformation method, horizontal scaling ratio, vertical scaling ratio, cropping start position, padding offset, original image size, output image size, and effective region mask.
[0244] The inverse spatial transformation parameters are those used to eliminate the effects of the forward spatial transformation in step S203 and restore the output image coordinates to the original image coordinate system. Let the scaling ratios in the horizontal and vertical directions be respectively... and The horizontal and vertical starting coordinates of the cropping window in the original image are respectively... and The horizontal and vertical fill offsets of the effective image region in the output image are respectively and Then the inverse space transformation can be expressed as: in, and They represent the first The x and y coordinates of the original image corresponding to each camera; and These represent the x and y coordinates of the output image, respectively. and These represent the horizontal scaling ratio and the vertical scaling ratio used for the forward spatial transformation, respectively. and These represent the horizontal and vertical starting coordinates of the cropping window in the original image, respectively. and These represent the horizontal and vertical fill offsets of the effective image region in the output image, respectively. The scaling ratio, cropping start coordinates, and fill offsets mentioned above together constitute the inverse spatial transformation parameters.
[0245] When the spatial transformation method is direct scaling, it can be set as follows: , , and All are 0. When the spatial transformation method is scaled fill, it can be set to 0. and The value is 0, and it is determined based on the starting position of the effective image region in the output image. and When the spatial transformation method includes cropping the region of interest or safe cropping, the starting position of the cropping window in the original image can be used to determine the cropping method. and .
[0246] If the spatial transformation metadata records the position of the cropping window in the scaled image, this position can be transformed to the original image coordinate system according to the corresponding scaling ratio before being determined. and When the forward spatial transformation also includes rotation, flipping, perspective transformation, or other affine transformations, a corresponding inverse transformation matrix can be established based on the spatial transformation metadata, and this inverse transformation matrix can be used to convert the output image coordinates into the original image coordinates.
[0247] After obtaining the original image coordinates, the valid original image coordinates are determined from the original image coordinates based on the spatial transformation metadata.
[0248] Valid original image coordinates refer to the original image coordinates whose corresponding output image coordinates lie within the valid image region and remain within the original image boundary range after the inverse transformation. Let... Indicates the first The effective region mask of the output image of each camera can be represented as: in, Indicates the first The set of effective potential spatial locations corresponding to each camera; Indicates the valid region mask; and They represent the first The original image width and original image height of each camera. For those belonging to The potential spatial location, obtained by inverse transformation and The coordinates were determined to be valid original image coordinates.
[0249] For potential spatial locations formed by padding pixels, the effective region mask value corresponding to their output image coordinates is 0; for locations that fall outside the boundaries of the original image after the inverse transform, their original image coordinates do not satisfy the above boundary conditions. These locations do not correspond to the true effective content in the original image and can be marked as invalid locations, and excluded using the validity mask during subsequent spatial location encoding or policy model processing.
[0250] Next, based on the spatial transformation method, the target mapping coordinates corresponding to each potential spatial location are determined from the output image coordinates and the effective original image coordinates.
[0251] Target mapping coordinates refer to the coordinates ultimately used for normalization to a unified reference grid; the width and height of the coordinate domain containing the target mapping coordinates are denoted as... and .
[0252] When the spatial transformation is a direct scaling or a size transformation that does not change the effective field of view, the output image coordinates can be used as the target mapping coordinates, and the output image width and height can be used as the target coordinate domain dimensions. When the spatial transformation involves cropping or padding, to preserve the actual position of the potential spatial location within the original camera's field of view, the effective original image coordinates can be used as the target mapping coordinates, and the original image width and height can be used as the target coordinate domain dimensions. The target mapping coordinates and their coordinate domain dimensions can be expressed as: The first line above applies to spatial transformation methods that directly scale or do not change the effective field of view. and These represent the output image width and height, respectively; the second row is applicable to spatial transformations that include cropping or padding. and The original image coordinates are valid.
[0253] In direct scaling, using the output image coordinates avoids unnecessary inverse spatial transformations, while ensuring that the complete boundary of the output image still corresponds to the complete field-of-view boundary of the original image. In cropping or padding, using valid original image coordinates corrects for positional changes caused by cropping start positions and padding offsets, allowing the target mapping coordinates to retain their original positional meaning within the camera's field of view.
[0254] After determining the target mapping coordinates, reference grid mapping is performed on each potential spatial location based on the target mapping coordinates and the grid size of the unified reference grid to determine the continuous coordinates of the reference grid corresponding to each potential spatial location.
[0255] Specifically, normalization can also be performed using the dimensions of the target coordinate domain where the target mapping coordinates lie. Let the continuous abscissa and continuous ordinate of the reference grid corresponding to the potential spatial location be respectively... and Then it can be expressed as: in, and They represent the first The continuous horizontal and vertical coordinates of the reference grid corresponding to each camera; and These represent the x-coordinate and y-coordinate of the target mapping, respectively. and These represent the width and height of the target coordinate domain, respectively. and These represent the width and height of the unified reference grid, respectively; max indicates taking the maximum value.
[0256] This formula maps the upper left and lower right boundaries of the target coordinate domain to the upper left and lower right boundaries of a unified reference grid, respectively. The same relative positions in different target coordinate domains can be mapped to continuous coordinates on the same or similar reference grid, thus eliminating the influence of different output image sizes and different original image sizes on the coordinate scale.
[0257] When the size of the target coordinate domain in a certain direction is 1, there are no multiple distinguishable positions in that direction. The "max" in the formula is used to avoid a denominator of 0 when the size of the target coordinate domain is 1; in this case, the continuous coordinates of the reference grid in that direction can be set to 0 or the intermediate coordinates of the corresponding direction of the unified reference grid.
[0258] Continuous coordinates of the reference grid can retain their decimal parts without needing to be rounded to discrete positions within the unified reference grid. For example, when a potential spatial location is mapped to between two adjacent integer positions on the unified reference grid, the decimal coordinates can be used to generate subsequent spatial rotation position codes to preserve the relative positional differences between different actual potential spatial grids.
[0259] When floating-point calculation errors cause the target mapping coordinates corresponding to a valid location to slightly exceed the boundary of the target coordinate domain, the location can be restricted to the boundary range of the target coordinate domain after confirming its validity. For invalid locations outside the filled region or the original image boundary, they should be excluded or marked separately based on the validity mask, rather than converting invalid locations into valid locations through boundary restrictions.
[0260] Using the above method, the same potential spatial location index in different clipping windows can be mapped to different reference grid continuous coordinates according to their respective clipping start positions, thereby preserving the difference of the position in the original camera field of view; the same original field of view position can be mapped to the same or similar reference grid continuous coordinates under different output resolutions or different clipping scales.
[0261] It should be noted that the unified reference grid is used to standardize the positional scale within the camera's field of view and the representation of positional coordinates under different input resolutions; it does not directly establish a 3D geometric correspondence between different cameras. When it is necessary to establish a geometric correspondence between different cameras, it can be further processed by combining camera intrinsic parameters, camera extrinsic parameters, depth information, homography transformation, or cross-view feature matching results.
[0262] Step S205 first determines the output image coordinates corresponding to the latent spatial location using the VAE spatial downsampling factor. Then, it uses spatial transformation metadata and inverse spatial transformation parameters to determine the effective original image coordinates and selects the target mapping coordinates based on the spatial transformation method. Thus, without changing the actual size of the latent features, it generates continuous reference grid coordinates with a uniform scale for latent spatial locations corresponding to different cameras and output resolutions, providing a coordinate basis for subsequent generation of spatial rotation position encoding.
[0263] Step S206: Based on the continuous coordinates of the reference grid corresponding to each potential spatial location, generate the spatial rotation position code corresponding to each potential spatial location, and use the spatial rotation position code to perform position encoding on the features at the corresponding potential spatial locations in each potential feature.
[0264] In this embodiment, spatial rotation position encoding refers to converting the two-dimensional spatial coordinates corresponding to the potential spatial location into a rotation phase, and then using the rotation phase to perform rotation transformation on the query vector and key vector in the policy model. Spatial rotation position encoding can be two-dimensional rotation position encoding, also known as spatial RoPE.
[0265] The continuous coordinates of the reference grid obtained in step S205 can retain the decimal part. Unlike the position lookup method that only accepts integer position indices, spatial rotation position encoding can directly calculate the rotation phase based on the continuous coordinates, thereby preserving the relative position differences between different actual potential spatial grids and avoiding interpolating potential features of different sizes to the same grid first.
[0266] refer to Figure 5 , Figure 5 This is an exploded flowchart of an embodiment of step S206 of this application.
[0267] like Figure 5As shown, in some optional embodiments, step S206 may include steps S2061 to S2064.
[0268] Step S2061: Obtain the rotation position encoding frequency sequence, rotation dimension, and spatial axis channel allocation method used by the strategy model.
[0269] The rotation position encoding frequency sequence specifies the angular frequency of different rotation channel pairs as they change with spatial coordinates. The rotation dimension specifies the number of channels in the query vector and key vector that participate in rotation position encoding. The spatial axis channel allocation method specifies which channel pairs participating in rotation are used to encode the height direction position and which are used to encode the width direction position.
[0270] Let the channel dimension of a single attention head in the policy model be . The rotation dimension is ,in Not greater than The positive even number. If the channels involved in the rotation are grouped into a rotation channel pair by pairing adjacent channels, then the number of rotation channel pairs L and the rotation position encoded frequency sequence can be expressed as: Where L represents the number of rotation channel pairs; Ω represents the rotation position encoded frequency sequence; This represents the angular frequency corresponding to the ℓth rotation channel pair.
[0271] The rotation position encoding frequency sequence can be directly read from the model configuration of the policy model, or it can be generated according to a preset frequency cardinality and rotation dimension. For example, different rotation channel pairs can be arranged in order from high frequency to low frequency or from low frequency to high frequency. Regardless of the generation method, the same frequency sequence, rotation dimension, channel pairing order, and spatial axis channel allocation method should be used in both the training and inference phases.
[0272] When the policy model includes multiple attention layers or multiple attention heads, each attention layer or attention head can share the same frequency sequence and spatial axis channel allocation method, or it can use pre-configured different frequency sequences or different rotation dimensions. The specific configuration should be consistent with the configuration used during the training of the policy model.
[0273] Step S2062: Based on the continuous coordinates of the reference grid corresponding to each potential spatial location and the rotation position coding frequency sequence, calculate the height direction rotation phase and width direction rotation phase corresponding to each potential spatial location, and generate the spatial rotation position code corresponding to each potential spatial location.
[0274] Let the first The potential spatial location corresponding to each camera is The continuous ordinate and continuous abscissa of the reference grid determined in step S205 are respectively and For the first frequency in the frequency sequence The rotational phase in the height direction and the rotational phase in the width direction corresponding to this potential spatial location, given the angular frequencies, can be expressed as follows: in, Indicates the first Potential spatial location of each camera In the Phase rotation in the height direction at each angular frequency; Indicates the corresponding rotation phase in the width direction; and These represent the continuous ordinates and continuous abscissas of the reference grid, respectively. This represents the angular frequency corresponding to the ℓth rotation channel pair.
[0275] because and The value can be a non-integer, and the rotation phase generated by this formula can also be a continuous value. Therefore, adjacent continuous coordinates in the same reference grid can produce a smoothly changing rotation phase without needing to round the continuous coordinates of the reference grid to discrete grid indices.
[0276] Spatial rotational position encoding can be composed of the sine and cosine values corresponding to the rotation phases in each height and width directions. For ease of explanation, let... This represents the rotation phase in any height direction or the rotation phase in any width direction, for a two-dimensional channel pair. The rotation operation can be defined as: in, Indicates the rotation phase as Two-dimensional rotation operation; a and b represent two channel components in a rotation channel pair; Represents the cosine function; This represents a sine function. The formula maintains the vector magnitude of the channel pair unchanged and injects spatial position information into the channel pair by rotating the angle.
[0277] The sine and cosine values of the required rotation phase can be pre-calculated for each potential spatial location to obtain the spatial rotation position code corresponding to that potential spatial location. For the same reference grid continuous coordinates, frequency sequence, and spatial axis channel allocation method, the corresponding sine and cosine values can be reused to reduce redundant calculations.
[0278] Step S2063: Generate query vectors and key vectors corresponding to each potential spatial location based on the features at the corresponding potential spatial locations in each potential feature.
[0279] Let the first Latent features of a camera in latent spatial location The eigenvector at that location is In the target attention layer of the policy model, the feature vector can be linearly projected using query projection parameters and key projection parameters to obtain the query vector and key vector: in, and Representing potential spatial locations The corresponding query vector and key vector; and These represent the query projection matrix and the key projection matrix, respectively. and These represent the optional query bias vector and key bias vector, respectively. This represents the feature vector at the corresponding potential spatial location in the latent features.
[0280] When the policy model employs multi-head attention, the projection results can be divided according to the attention heads, and subsequent rotation transformations can be performed on the query vector and key vector in each attention head. The query projection parameters and key projection parameters are policy model parameters, which can be learned during model training and do not need to be re-determined due to changes in the current output image resolution.
[0281] Step S2064: Based on the spatial rotation position encoding, rotation dimension, and spatial axis channel allocation method corresponding to each potential spatial position, perform rotation transformation on the query vector and key vector corresponding to each potential spatial position, and determine the rotated query vector and key vector as the position encoding result of the feature at the corresponding potential spatial position in each potential feature.
[0282] Specifically, it can be Each rotational channel pair is divided into a set of channel pairs in the height direction. and width direction channel pair set The two sets do not overlap and together cover all channel pairs involved in the rotation: in, This represents the set of rotational channel pairs assigned to the height direction; This represents the set of rotation channel pairs allocated to the width direction. The spatial axis channels can be allocated in an even manner, so that the height and width directions use the same or similar number of rotation channel pairs; or a non-uniform allocation method can be used depending on the sensitivity of the task to the vertical or horizontal position.
[0283] For belonging to The Each rotating channel pair uses a phase rotation in the height direction; for those belonging to The A pair of rotating channels, using a phase rotation in the width direction. Let... Indicates the first Each rotating channel corresponds to a spatial axis. exist belong Time to take ,exist belong Time to take Then the rotation result of the corresponding rotation channel pairs of the query vector and key vector can be expressed as: in, and Represents the query vector number 1 Two channel components in a rotating channel pair; and This represents the two corresponding channel components in the key vector; and These represent the rotation results of the corresponding query vector rotation channel pair and key vector rotation channel pair, respectively; R represents a two-dimensional rotation operation. This indicates the height-direction rotation phase or the width-direction rotation phase selected based on the spatial axis channel allocation method.
[0284] The two expressions above use the same spatial rotation position encoding for the query vector and key vector at the same potential spatial location. Therefore, when subsequently calculating the attention correlation between the query vector and the key vector, the relative rotation phase between different potential spatial locations can reflect the relative positional relationship between their continuous coordinates on the reference grid.
[0285] For channels in the query vector and key vector that are not included in the rotation dimension, their original channel values can be preserved. Let... This indicates the channel number in the attention head. The channel numbers that did not participate in the rotation satisfy the following: The above channel range only performs spatial rotation position encoding on channels within the rotation dimension; channels outside the rotation dimension retain their original content features. The rotation dimension can be equal to or less than the attention head channel dimension.
[0286] The rotated query vector and key vector are used to determine the positional encoding results of features at the corresponding latent spatial locations. Subsequent attention operations can directly use the rotated query vector and key vector to calculate the attention score, while the value vector can remain unchanged. When the policy model adopts other preset positional encoding structures, the value vector or other intermediate features can also be processed accordingly according to that structure.
[0287] For potential spatial locations marked as invalid in step S205, a spatial rotation position code can be omitted, or a placeholder position code can be generated and a validity mask can be used to prevent that location from participating in the effective attention calculation. This avoids interference from filled regions or locations outside the original image boundaries with the policy model.
[0288] The spatial rotation position code generated in step S206 is used to represent the two-dimensional position in the unified reference grid; camera identity, temporal position, or motion phase can be represented by additional codes respectively.
[0289] Step S206 allows the use of continuous coordinates from a reference grid to inject uniform-scale two-dimensional positional information into potential spatial locations within potential features of different sizes, without requiring spatial interpolation. By performing rotation transformations on the query vector and key vector, the attention computation becomes capable of perceiving the relative spatial relationships between potential spatial locations, thereby supporting the policy model in processing actual potential spatial grids formed by different camera output resolutions.
[0290] This application first obtains the video frames acquired by each camera in the current multi-camera setup, the total pixel budget corresponding to the current multi-camera setup, the VAE spatial downsampling factor, and the unified reference grid. Second, based on the total pixel budget and the VAE spatial downsampling factor, it determines the multi-camera output resolution configuration of the current multi-camera setup. The multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera setup, where the height and width of each output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget. Third, based on the multi-camera output resolution configuration, it performs a spatial transformation on each video frame to generate the corresponding VAE output. The system first takes an image as input and records the spatial transformation metadata corresponding to the spatial transformation. Then, it performs VAE encoding on each VAE input image to generate latent features and determines the actual latent spatial grid corresponding to each latent feature. Further, based on the unified reference grid, each actual latent spatial grid, and the spatial transformation metadata, it performs reference grid mapping on the latent spatial positions in each actual latent spatial grid to determine the continuous coordinates of the reference grid corresponding to each latent spatial position. Finally, based on the continuous coordinates of the reference grid corresponding to each latent spatial position, it generates the spatial rotation position code corresponding to each latent spatial position and uses the spatial rotation position code to perform position encoding on the features at the corresponding latent spatial positions in each latent feature.
[0291] Through the above technical solution, this application uniformly determines the output image size of each camera in the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor, ensuring that the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget, and that the height and width of each output image size are integer multiples of the VAE spatial downsampling factor. Therefore, it is possible to control the total number of pixels used in multi-camera video processing while adapting the output image size of each camera to the VAE spatial downsampling process, thereby simultaneously satisfying both the multi-camera total pixel budget constraint and the VAE spatial downsampling constraint, and improving the controllability of computing resources.
[0292] Simultaneously, this application records the corresponding spatial transformation metadata when performing spatial transformation on each video frame, and performs reference grid mapping on the potential spatial positions in each actual potential spatial grid by combining a unified reference grid and each actual potential spatial grid. Therefore, it is possible to consider the positional changes caused by different output image sizes and different spatial transformation methods during the reference grid mapping process, and uniformly represent the potential spatial positions in different actual potential spatial grids as continuous coordinates in the same reference coordinate system.
[0293] Based on this, this application generates spatial rotation position codes according to the continuous coordinates of the reference grid corresponding to each potential spatial location, so that locations with the same or similar spatial semantics have comparable position coding phases under different actual potential spatial grids and different spatial transformation conditions. This reduces the impact of changes in the size of the actual potential spatial grid and video frame spatial transformations on the position coding scale, maintains the consistency of the position coding scale under different actual potential spatial grids and different spatial transformation conditions, improves the consistency and stability of the spatial position representation of multi-camera potential features, and is beneficial for subsequent strategy models to uniformly process multi-camera potential features.
[0294] refer to Figure 6 , Figure 6 A schematic diagram of an embodiment of a multi-camera video resolution configuration and spatial location encoding apparatus according to this application is shown.
[0295] like Figure 6 As shown, the resolution configuration and spatial location coding device 600 for multi-camera video includes an information acquisition unit 601, a resolution configuration determination unit 602, a spatial transformation unit 603, a VAE coding unit 604, a reference grid mapping unit 605, and a location coding unit 606.
[0296] The information acquisition unit 601 is used to acquire video frames captured by each camera in the current multi-camera combination, the total pixel budget corresponding to the current multi-camera combination, the VAE spatial downsampling factor, and the unified reference grid.
[0297] The resolution configuration determination unit 602 is used to determine the multi-camera output resolution configuration of the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor. The multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera combination, where the height and width of each output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget.
[0298] The spatial transformation unit 603 is used to perform spatial transformation on each video frame according to the multi-camera output resolution configuration, generate the corresponding VAE input image, and record the spatial transformation metadata corresponding to the spatial transformation.
[0299] VAE encoding unit 604 is used to perform VAE encoding on each VAE input image, generate latent features, and determine the actual latent space grid corresponding to each latent feature.
[0300] The reference grid mapping unit 605 is used to perform reference grid mapping on the potential spatial locations in each actual potential spatial grid based on the unified reference grid, each actual potential spatial grid and spatial transformation metadata, and to determine the continuous reference grid coordinates corresponding to each potential spatial location.
[0301] The position encoding unit 606 is used to generate a spatial rotation position code corresponding to each potential spatial position based on the continuous coordinates of the reference grid corresponding to each potential spatial position, and to use the spatial rotation position code to perform position encoding on the features at the corresponding potential spatial positions in each potential feature.
[0302] In some optional implementations, the resolution configuration determination unit 602 is further configured to: Obtain a pre-established set of finitely valid multi-camera discrete resolution configurations corresponding to the current multi-camera combination and the total pixel budget. The set of finitely valid multi-camera discrete resolution configurations includes multiple sets of candidate multi-camera resolution configurations. Each set of candidate multi-camera resolution configurations includes the candidate output image size of each camera in the current multi-camera combination. The height and width of each candidate output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each candidate output image size does not exceed the total pixel budget.
[0303] Based on the total pixel budget and VAE spatial downsampling factor, determine the multi-camera output resolution configuration of the current multi-camera combination from a finite set of legal multi-camera discrete resolution configurations.
[0304] In some optional implementations, the resolution configuration determination unit 602 is further configured to: Obtain the horizontal effective field of view, vertical effective field of view, and preset minimum angle sampling density of each camera in the current multi-camera combination.
[0305] The base output image size of each camera is determined based on the effective horizontal field of view, the effective vertical field of view, and the preset minimum angle sampling density.
[0306] Based on the basic output image size, total pixel budget, and VAE spatial downsampling factor, determine the multi-camera output resolution configuration of the current multi-camera combination from a finite set of legal multi-camera discrete resolution configurations.
[0307] In some optional implementations, the resolution configuration determination unit 602 is further configured to: Obtain at least one piece of task information from the following: current action phase, camera role of each camera, area of key task region, and visibility of key content.
[0308] Based on the mission information, determine the dynamic budget requirements for each camera in the current multi-camera setup.
[0309] Based on the dynamic budget requirements and the total pixel budget, the continuous target pixel budget for each camera is determined, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget.
[0310] Based on the pixel budget of each continuous target and the VAE spatial downsampling factor, the multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations.
[0311] In some optional implementations, the resolution configuration determination unit 602 is further configured to: Obtain task features and low-cost camera content features for each camera in the current multi-camera setup.
[0312] Input the task features, content features of each low-cost camera, and a finite set of legal multi-camera discrete resolution configurations into the learnable configuration selection model to generate configuration selection scores for each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations.
[0313] Based on the scores selected for each configuration, the multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations.
[0314] In some optional implementations, the resolution configuration determination unit 602 is further configured to: Obtain the original image dimensions of each camera in the current multi-camera setup.
[0315] Based on the total pixel budget, determine the continuous target pixel budget for each camera, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget.
[0316] Calculate the continuous ideal output image size for each camera based on the original aspect ratio corresponding to each original image size and the pixel budget for each continuous target.
[0317] Based on the size of each continuous ideal output image, calculate the pixel budget error and aspect ratio error corresponding to each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations.
[0318] Based on pixel budget error and aspect ratio error, the multi-camera output resolution configuration of the current multi-camera combination is jointly determined from a finite set of legal multi-camera discrete resolution configurations.
[0319] In some alternative implementations, the spatial transformation unit 603 is further configured to: Obtain the original image size, full field-of-view preservation requirements, and region of interest information for each video frame.
[0320] Calculate the corresponding aspect ratio differences based on the original image size and the output image size in the multi-camera output resolution configuration.
[0321] The reliability of the region of interest is determined based on the region of interest information.
[0322] Based on the aspect ratio difference, the need to preserve the entire field of view, and the reliability of the region of interest, a spatial transformation method is selected. The spatial transformation methods include direct scaling, scaling and then filling, scaling and then safely cropping, or cropping the region of interest.
[0323] Perform spatial transformation on each video frame according to the spatial transformation method to generate the corresponding VAE input image.
[0324] Record the spatial transformation parameters corresponding to the spatial transformation method and generate spatial transformation metadata.
[0325] In some alternative implementations, the spatial transformation unit 603 is further configured to: Obtain historical operation region of interest information corresponding to the current action stage and previous video frames, and determine the region center, region scale, and region confidence of the current video frame from the operation region of interest information.
[0326] Based on the current action stage and historical operation information about the region of interest, time smoothing is performed on the region center and the region scale to determine candidate clipping windows.
[0327] Verify the coverage of the candidate clipping window with the key content of the task and the preset safety boundaries.
[0328] The reliability of the region of interest is determined based on the regional confidence level and coverage.
[0329] In some alternative implementations, the spatial transformation unit 603 is further configured to: Compare the aspect ratio difference with the preset direct scaling threshold.
[0330] Verify whether the safe cropping after scaling preserves the critical content of the task.
[0331] If the aspect ratio difference does not exceed the preset direct scaling threshold, direct scaling is selected.
[0332] In response to aspect ratio differences exceeding the preset direct scaling threshold and the need to preserve the full field of view, select to scale and then fill.
[0333] In response to situations where the aspect ratio difference exceeds a preset direct scaling threshold, there is no need to retain the entire field of view, and the reliability of the region of interest meets the preset reliability conditions, the region of interest is selected for cropping.
[0334] In response to situations where the aspect ratio difference exceeds the preset direct scaling threshold, there is no need to retain the complete field of view, the reliability of the region of interest does not meet the preset reliability conditions, and scaling is safe for cropping to retain the key content of the task, select scaling for safe cropping.
[0335] In response to situations where the aspect ratio difference exceeds the preset direct scaling threshold, there is no need to retain the complete field of view, the reliability of the area of interest does not meet the preset reliability conditions, and the scaling and safe cropping does not retain the task's critical content, the option to roll back and fill after scaling is selected.
[0336] In some alternative implementations, the reference mesh mapping unit 605 is further configured to: Based on the VAE spatial downsampling factor and each actual potential space grid, determine the output image coordinates corresponding to the potential spatial location in each actual potential space grid.
[0337] Based on the spatial transformation metadata, determine the inverse spatial transformation parameters corresponding to the spatial transformation method, convert each output image coordinate into the original image coordinates based on the inverse spatial transformation parameters, and determine the valid original image coordinates from the original image coordinates based on the spatial transformation metadata.
[0338] Based on the spatial transformation method, the target mapping coordinates corresponding to each potential spatial location are determined from the output image coordinates and the effective original image coordinates. Then, based on the target mapping coordinates and the grid size of the unified reference grid, reference grid mapping is performed on each potential spatial location to determine the continuous coordinates of the reference grid corresponding to each potential spatial location.
[0339] In some alternative implementations, the reference mesh mapping unit 605 is further configured to: In response to the spatial transformation method being direct scaling, the target mapping coordinates corresponding to each potential spatial location are determined based on each actual potential spatial grid and each output image coordinate.
[0340] In response to any of the spatial transformation methods of scaling and filling, scaling and safe cropping, or operation of region of interest cropping, without preserving the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial location are determined based on each actual potential spatial grid and each output image coordinate.
[0341] In response to any of the spatial transformation methods of scaling and filling, scaling and safe cropping, or operation of region of interest cropping, and in order to preserve the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial location are determined based on each valid original image coordinate and each original image size.
[0342] In some alternative implementations, the position encoding unit 606 is further configured to: The strategy model is obtained by analyzing the rotation position encoding frequency sequence, rotation dimension, and spatial axis channel allocation method.
[0343] Based on the continuous coordinates of the reference grid corresponding to each potential spatial location and the rotation position coding frequency sequence, the height direction rotation phase and width direction rotation phase corresponding to each potential spatial location are calculated, and the spatial rotation position coding corresponding to each potential spatial location is generated.
[0344] Based on the features at the corresponding potential spatial locations in each potential feature, generate the query vector and key vector corresponding to each potential spatial location.
[0345] Based on the spatial rotation position encoding, rotation dimension, and spatial axis channel allocation method corresponding to each potential spatial location, a rotation transformation is performed on the query vector and key vector corresponding to each potential spatial location, and the rotated query vector and key vector are determined as the position encoding result of the feature at the corresponding potential spatial location in each potential feature.
[0346] In some optional implementations, the resolution configuration determination unit 602 is further configured to: Obtain the current task requirements, the current action stage, the action stage of the previous processing time, and the historical multi-camera output resolution configuration used in the previous processing time.
[0347] Based on the total pixel budget, VAE spatial downsampling factor, and current task requirements, determine the configuration evaluation value of each candidate multi-camera resolution configuration in the finite set of legal multi-camera discrete resolution configurations.
[0348] Candidate multi-camera output resolution configurations are determined based on the evaluation values of each configuration.
[0349] In response to the existence of historical multi-camera output resolution configurations, calculate the configuration gain of the candidate multi-camera output resolution configuration relative to the historical multi-camera output resolution configuration.
[0350] In response to the absence of a historical multi-camera output resolution configuration, a candidate multi-camera output resolution configuration is determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0351] In response to the existence of historical multi-camera output resolution configurations, the current action stage being the same as the action stage at the previous processing moment, and the configuration gain not exceeding a preset switching threshold, the historical multi-camera output resolution configuration is determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0352] In response to the existence of historical multi-camera output resolution configurations and the fact that the current action phase is different from the action phase at the previous processing time, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination.
[0353] In response to the existence of historical multi-camera output resolution configurations, the current action phase being the same as the action phase at the previous processing moment, and the configuration benefit exceeding a preset switching threshold, the candidate multi-camera output resolution configuration is determined as the multi-camera output resolution configuration of the current multi-camera combination.
[0354] The specific processing of each module, the meaning of related terms, optional implementation methods, and their technical effects in the device embodiments can be found in the descriptions of the corresponding steps in the foregoing method embodiments, and will not be repeated here. The above modules can be implemented by software, hardware, firmware, or a combination thereof, and can also be combined or further split according to the deployment method of computing and storage resources.
[0355] The following is for reference. Figure 7 It shows a schematic diagram of the structure of a computer system 700 suitable for implementing the electronic device of the present application. Figure 7 The computer system 700 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.
[0356] In some alternative embodiments, this application also provides an electronic device including one or more processors and a storage device storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement any of the foregoing method embodiments.
[0357] Specifically, one or more processors can correspond to Figure 7 The processing device 701 shown may include a central processing unit, a graphics processing unit, an artificial intelligence accelerator, a tensor processor, or other general-purpose or special-purpose processors. Different processors may respectively perform at least one of the following processing steps: multi-camera video frame acquisition, multi-camera output resolution configuration determination, video frame spatial transformation, VAE encoding, reference mesh mapping, and spatial rotation position encoding.
[0358] The storage device may include at least one of read-only memory (ROM) 702, random access memory (RAM) 703, and storage device 708. ROM 702 may store a boot program or a firmware program; RAM 703 may store programs and data used during the operation of the electronic device; storage device 708 may include a disk, solid-state drive, flash memory device, or other non-volatile storage device, and may store one or more programs for implementing resolution configuration and spatial location encoding methods for multi-camera video.
[0359] Specifically, one or more programs may include program instructions for acquiring video frames, total pixel budget, VAE spatial downsampling factor, and uniform reference grid for each camera in the current multi-camera combination; program instructions for determining the multi-camera output resolution configuration based on the total pixel budget and VAE spatial downsampling factor; program instructions for performing spatial transformation on each video frame according to the multi-camera output resolution configuration and recording spatial transformation metadata; and program instructions for performing VAE encoding on the spatially transformed VAE input image and generating latent features.
[0360] One or more programs may also include program instructions for performing reference grid mapping on each potential spatial location based on a unified reference grid, the actual potential spatial grid corresponding to each potential feature, and spatial transformation metadata, as well as program instructions for generating spatial rotation position codes based on the continuous coordinates of the reference grid corresponding to each potential spatial location, and using the spatial rotation position codes to position-encode the features at the corresponding potential spatial locations in each potential feature.
[0361] One or more programs may further include program instructions for obtaining a finite set of legal multi-camera discrete resolution configurations, program instructions for selecting a multi-camera output resolution configuration based on at least one of the following: effective field of view, mission information, low-cost camera content features, continuous target pixel budget, or original aspect ratio, and program instructions for controlling the switching of multi-camera output resolution configurations based on the current action phase, historical multi-camera output resolution configurations, and configuration benefits.
[0362] One or more programs may also include program instructions for determining the reliability of the region of interest, program instructions for selecting a spatial transformation method from direct scaling, scale-fill, scale-safe cropping, and operational region of interest cropping, program instructions for determining inverse spatial transformation parameters and valid original image coordinates based on spatial transformation metadata, and program instructions for determining target mapping coordinates based on whether the positional relationships in the original camera field of view need to be preserved.
[0363] One or more programs may further include program instructions for obtaining the rotation position encoding frequency sequence, rotation dimension, and spatial axis channel allocation method; program instructions for calculating the rotation phase in the height direction and the rotation phase in the width direction; program instructions for generating query vectors and key vectors based on the features at each potential spatial location; and program instructions for performing rotation transformations on the query vectors and key vectors using spatial rotation position encoding.
[0364] like Figure 7As shown, the processing device 701, ROM 702, and RAM 703 can be connected to each other via bus 704. I / O interface 705 can also be connected to bus 704. When the electronic device is running, the processing device 701 can perform processing corresponding to the resolution configuration and spatial location encoding method of multi-camera video according to the program stored in ROM 702 or according to the program loaded from storage device 708 into RAM 703.
[0365] I / O interface 705 can connect to input device 706, output device 707, storage device 708, and communication device 709. Input device 706 may include a keyboard, mouse, touch device, camera control device, or other device for inputting device operating parameters. Output device 707 may include a display, status indicator, or other device for outputting multi-camera output resolution configuration, spatial transformation results, device operating status, and model processing results.
[0366] The communication device 709 enables wired or wireless communication between electronic devices and multi-camera acquisition devices, robot task control devices, model training nodes, model inference nodes, configuration storage nodes, or other electronic devices. Electronic devices can receive video frames acquired by each camera, current task requirements, current action stage, camera role, region of interest information, and device status information through the communication device 709, and can provide location-encoded latent features to the model training nodes or model inference nodes.
[0367] In some embodiments, the processing device 701 may include multiple processors. One processor may perform video frame acquisition, resolution configuration determination, and spatial transformation, while another processor may perform VAE encoding, reference grid mapping, and spatial rotation position encoding. The multiple processors may share RAM 703 or each have their own device memory, and may transmit VAE input images, latent features, spatial transformation metadata, reference grid continuous coordinates, or position encoding results via high-speed interconnects.
[0368] In some implementations, the electronic device can be a robot control device, a model training server, a model inference server, an edge computing device, or other computing devices with image processing capabilities, or it can be a distributed electronic device composed of multiple computing nodes. When the electronic device is a distributed electronic device, the information acquisition unit 601, the resolution configuration determination unit 602, the spatial transformation unit 603, the VAE encoding unit 604, the reference mesh mapping unit 605, and the position encoding unit 606 can be deployed on the same computing node, or they can be deployed on different computing nodes that can communicate with each other.
[0369] Components in electronic devices are not required to all conform to Figure 7The setup is as shown. Depending on the specific implementation requirements, the electronic device may omit some input or output devices, or it may add a graphics processor, artificial intelligence accelerator, direct memory access controller, camera access interface, high-speed network interface, or other components suitable for performing multi-camera video processing and strategy model calculations.
[0370] In some alternative embodiments, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the aforementioned method embodiments.
[0371] Specifically, a computer-readable storage medium can be a tangible medium capable of containing or storing computer programs. Computer-readable storage media can include portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, solid-state drives, optical storage devices, magnetic storage devices, or any combination of the above media.
[0372] The computer-readable storage medium may be disposed within the aforementioned electronic device or independently of the aforementioned electronic device. The computer program stored in the computer-readable storage medium may include program instructions for causing the processor to perform multi-camera output resolution configuration determination, video frame spatial transformation, VAE encoding, reference grid mapping, and spatial rotation position encoding, and may also include program instructions for performing operational region of interest reliability determination, spatial transformation mode selection, multi-camera output resolution configuration switching, and invalid position masking processing.
[0373] When a computer program is executed by one or more processors, the one or more processors may perform all or some of the steps in the foregoing method embodiments. When executed jointly by multiple processors, different processors may respectively perform different processes in video frame spatial transformation, VAE encoding, reference mesh mapping, and spatial rotation position encoding.
[0374] In some alternative implementations, this application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement any of the foregoing method embodiments.
[0375] Specifically, the computer program product may include a computer program or computer instructions for implementing a resolution configuration and spatial location coding method for multi-camera video. The computer program or computer instructions may be downloaded from a network and installed on an electronic device, or loaded onto an electronic device from a computer-readable storage medium. When executed by a processor, the computer program or computer instructions cause the electronic device to determine a multi-camera output resolution configuration that satisfies the total pixel budget constraint and the VAE spatial downsampling constraint, process the video frames acquired by each camera according to the multi-camera output resolution configuration, and generate latent feature location coding results that preserve relative spatial relationships.
[0376] Computer programs or instructions can be written in one or more programming languages and can be executed entirely on a single electronic device, or separately on multi-camera acquisition nodes, image preprocessing nodes, VAE encoding nodes, policy model processing nodes, and robot control nodes. The program portions executed at different nodes can exchange video frames, multi-camera output resolution configurations, spatial transformation metadata, latent features, continuous coordinates of the reference grid, spatial rotation position encoding, and model processing results via wired or wireless networks.
[0377] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should be noted that in some alternative implementations, the functions marked in the blocks may be executed in a different order than that shown in the drawings. For example, two consecutively shown blocks may actually be executed substantially in parallel, or in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware system performing the specified function or operation, or by a combination of dedicated hardware and computer instructions.
[0378] The units involved in the embodiments of this application can be implemented in software, in hardware, or in a combination of both. The name of a unit does not necessarily limit the unit itself. For example, the reference grid mapping unit 605 can also be described as "a unit that determines the continuous coordinates of the reference grid based on a unified reference grid, the actual potential spatial grid, and spatial transformation metadata," and the position encoding unit 606 can also be described as "a unit that generates spatial rotation position codes based on the continuous coordinates of the reference grid and performs position encoding on potential features."
[0379] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with, but not limited to, technical features disclosed in this application that have similar functions.
Claims
1. A method for resolution configuration and spatial location coding of multi-camera video, characterized in that, The method includes: Obtain the video frames captured by each camera in the current multi-camera combination, the total pixel budget corresponding to the current multi-camera combination, the VAE spatial downsampling factor, and the unified reference grid; Based on the total pixel budget and the VAE spatial downsampling factor, the multi-camera output resolution configuration of the current multi-camera combination is determined, wherein the multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera combination, the height and width of each output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget; Based on the multi-camera output resolution configuration, spatial transformation is performed on each video frame to generate a corresponding VAE input image, and the spatial transformation metadata corresponding to the spatial transformation is recorded. VAE encoding is performed on each of the VAE input images to generate latent features, and the actual latent space grid corresponding to each latent feature is determined; Based on the unified reference grid, each of the actual potential space grids and the spatial transformation metadata, reference grid mapping is performed on the potential spatial locations in each of the actual potential space grids to determine the continuous reference grid coordinates corresponding to each of the potential spatial locations. Based on the continuous coordinates of the reference grid corresponding to each potential spatial location, a spatial rotation position code is generated for each potential spatial location, and the spatial rotation position code is used to perform position encoding on the features at the corresponding potential spatial locations in each potential feature.
2. The method according to claim 1, characterized in that, The step of determining the multi-camera output resolution configuration of the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain a pre-established set of finitely valid multi-camera discrete resolution configurations corresponding to the current multi-camera combination and the total pixel budget. The set of finitely valid multi-camera discrete resolution configurations includes multiple sets of candidate multi-camera resolution configurations. Each set of candidate multi-camera resolution configurations includes the candidate output image size of each camera in the current multi-camera combination. The height and width of each candidate output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each candidate output image size does not exceed the total pixel budget. The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor.
3. The method according to claim 2, characterized in that, The step of determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain the horizontal effective field of view, vertical effective field of view, and preset minimum angle sampling density of each camera in the current multi-camera combination; The basic output image size of each camera is determined based on the horizontal effective field of view, the vertical effective field of view, and the preset minimum angle sampling density. The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the respective base output image sizes, the total pixel budget, and the VAE spatial downsampling factor.
4. The method according to claim 2, characterized in that, The step of determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain at least one piece of task information from the current action phase, camera roles of each camera, area of the key task region, and visibility of key content; Based on the task information, determine the dynamic budget requirements of each camera in the current multi-camera setup; Based on the dynamic budget requirements and the total pixel budget, the continuous target pixel budget for each camera is determined, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget; The multi-camera output resolution configuration of the current multi-camera combination is determined from the finite set of legal multi-camera discrete resolution configurations based on the continuous target pixel budget and the VAE spatial downsampling factor.
5. The method according to claim 2, characterized in that, The step of determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Acquire task features and low-cost camera content features corresponding to each camera in the current multi-camera combination; The task features, the content features of each low-cost camera, and the set of finitely legal multi-camera discrete resolution configurations are input into a learnable configuration selection model to generate a configuration selection score for each candidate multi-camera resolution configuration in the set of finitely legal multi-camera discrete resolution configurations. Based on the selected scores for each configuration, the multi-camera output resolution configuration of the current multi-camera combination is determined from the set of finite legal multi-camera discrete resolution configurations.
6. The method according to claim 2, characterized in that, The step of determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain the original image size of each camera in the current multi-camera setup; Based on the total pixel budget, the continuous target pixel budget for each camera is determined, wherein the sum of the continuous target pixel budgets does not exceed the total pixel budget; Calculate the continuous ideal output image size of each camera based on the original aspect ratio corresponding to each original image size and the continuous target pixel budget of each. Based on the continuous ideal output image size, calculate the pixel budget error and aspect ratio error corresponding to each candidate multi-camera resolution configuration in the finite legal multi-camera discrete resolution configuration set; Based on the pixel budget error and the aspect ratio error, the multi-camera output resolution configuration of the current multi-camera combination is jointly determined from the finite set of legal multi-camera discrete resolution configurations.
7. The method according to claim 1, characterized in that, The step involves performing spatial transformation on each video frame according to the multi-camera output resolution configuration, generating a corresponding VAE input image, and recording the spatial transformation metadata corresponding to the spatial transformation, including: Obtain the original image size, full field-of-view preservation requirements, and region of interest information for each video frame; Calculate the corresponding aspect ratio difference based on the original image size and the output image size in the multi-camera output resolution configuration; The reliability of the region of interest is determined based on the region of interest information. Based on the aspect ratio difference, the requirement to preserve the complete field of view, and the reliability of the region of interest, a spatial transformation method is selected, wherein the spatial transformation method includes direct scaling, scaling and then filling, scaling and then safely cropping, or cropping of the region of interest. Perform spatial transformation on each of the video frames according to the spatial transformation method to generate the corresponding VAE input image; Record the spatial transformation parameters corresponding to the spatial transformation method, and generate the spatial transformation metadata.
8. The method according to claim 7, characterized in that, The step of determining the reliability of the region of interest based on the region of interest information includes: Obtain historical operation region of interest information corresponding to the current action stage and the preceding video frame, and determine the region center, region scale and region confidence corresponding to the current video frame from the operation region of interest information; Based on the current action stage and the region of interest information of the historical operation, time smoothing is performed on the region center and the region scale to determine candidate clipping windows; Verify the coverage of the candidate cropping window with the key content of the task and the preset security boundary; The reliability of the region of interest is determined based on the region confidence level and the coverage.
9. The method according to claim 7, characterized in that, The step of selecting a spatial transformation method based on the aspect ratio difference, the requirement to preserve the complete field of view, and the reliability of the region of interest includes: Compare the aspect ratio difference with a preset direct scaling threshold; Verify whether the scaled and safe cropping retains the key content of the task; In response to the aspect ratio difference not exceeding the preset direct scaling threshold, direct scaling is selected; In response to the aspect ratio difference exceeding the preset direct scaling threshold and the need to retain the complete field of view, the scaling and filling option is selected. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, and the reliability of the region of interest meeting the preset reliability condition, the region of interest is selected for cropping. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, the reliability of the region of interest not meeting the preset reliability condition, and the safe cropping after scaling retaining the key content of the task, the safe cropping after scaling is selected. In response to the aspect ratio difference exceeding the preset direct scaling threshold, the need to retain the complete field of view, the reliability of the region of interest not meeting the preset reliability condition, and the scaling-after safe cropping not retaining the key content of the task, the option to revert to scaling-after filling is selected.
10. The method according to claim 7, characterized in that, The step of performing reference grid mapping on potential spatial locations in each of the actual potential spatial grids based on the unified reference grid, each of the actual potential spatial grids, and the spatial transformation metadata, to determine the continuous reference grid coordinates corresponding to each potential spatial location, includes: Based on the VAE spatial downsampling factor and each of the actual latent space grids, determine the output image coordinates corresponding to the latent spatial positions in each of the actual latent space grids; Based on the spatial transformation metadata, determine the inverse spatial transformation parameters corresponding to the spatial transformation method, convert each of the output image coordinates into original image coordinates based on the inverse spatial transformation parameters, and determine the valid original image coordinates from the original image coordinates based on the spatial transformation metadata. According to the spatial transformation method, the target mapping coordinates corresponding to each potential spatial location are determined from the output image coordinates and the effective original image coordinates. Based on the target mapping coordinates and the grid size of the unified reference grid, reference grid mapping is performed on each potential spatial location to determine the continuous reference grid coordinates corresponding to each potential spatial location.
11. The method according to claim 10, characterized in that, Determining the target mapping coordinates corresponding to each potential spatial location from the output image coordinates and the effective original image coordinates according to the spatial transformation method includes: In response to the spatial transformation method being direct scaling, the target mapping coordinates corresponding to each potential spatial location are determined based on each actual potential spatial grid and each output image coordinate. In response to the spatial transformation method being any one of the scaling-up fill, scaling-up safe crop, or operation region of interest crop, and without needing to preserve the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial position are determined based on each actual potential spatial grid and each output image coordinate; In response to the spatial transformation method being any one of the scaling-up fill, scaling-up safe crop, or operation region of interest crop, and requiring the preservation of the positional relationships in the original camera field of view, the target mapping coordinates corresponding to each potential spatial location are determined based on each of the effective original image coordinates and each of the original image sizes.
12. The method according to claim 11, characterized in that, The step of generating a spatial rotation position code corresponding to each potential spatial position based on the continuous coordinates of the reference grid corresponding to each potential spatial position, and using the spatial rotation position code to perform position encoding on the features at the corresponding potential spatial positions in each potential feature, includes: The rotation position encoding frequency sequence, rotation dimension, and spatial axis channel allocation method used in the strategy model are obtained. Based on the continuous coordinates of the reference grid corresponding to each potential spatial location and the rotation position coding frequency sequence, calculate the height direction rotation phase and width direction rotation phase corresponding to each potential spatial location, and generate the spatial rotation position code corresponding to each potential spatial location; Based on the features at the corresponding potential spatial locations in each of the potential features, generate query vectors and key vectors corresponding to each potential spatial location; Based on the spatial rotation position encoding corresponding to each potential spatial location, the rotation dimension, and the spatial axis channel allocation method, a rotation transformation is performed on the query vector and key vector corresponding to each potential spatial location, and the rotated query vector and key vector are determined as the position encoding result of the feature at the corresponding potential spatial location in each potential feature.
13. The method according to claim 2, characterized in that, The step of determining the multi-camera output resolution configuration of the current multi-camera combination from the finite set of legal multi-camera discrete resolution configurations based on the total pixel budget and the VAE spatial downsampling factor includes: Obtain the current task requirements, the current action stage, the action stage of the previous processing moment, and the historical multi-camera output resolution configuration used in the previous processing moment; Based on the total pixel budget, the VAE spatial downsampling factor, and the current task requirements, determine the configuration evaluation value of each candidate multi-camera resolution configuration in the finite legal multi-camera discrete resolution configuration set; Candidate multi-camera output resolution configurations are determined based on the evaluation values of each configuration. In response to the existence of the historical multi-camera output resolution configuration, calculate the configuration gain of the candidate multi-camera output resolution configuration relative to the historical multi-camera output resolution configuration; In response to the absence of the historical multi-camera output resolution configuration, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration, the current action stage being the same as the action stage at the previous processing moment, and the configuration benefit not exceeding a preset switching threshold, the historical multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration and the fact that the current action phase is different from the action phase at the previous processing time, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination. In response to the existence of the historical multi-camera output resolution configuration, the current action stage being the same as the action stage at the previous processing moment, and the configuration benefit exceeding the preset switching threshold, the candidate multi-camera output resolution configuration is determined to be the multi-camera output resolution configuration of the current multi-camera combination.
14. A resolution configuration and spatial location encoding device for multi-camera video, characterized in that, include: The information acquisition unit is used to acquire video frames captured by each camera in the current multi-camera combination, the total pixel budget corresponding to the current multi-camera combination, the VAE spatial downsampling factor, and the unified reference grid. The resolution configuration determination unit is used to determine the multi-camera output resolution configuration of the current multi-camera combination based on the total pixel budget and the VAE spatial downsampling factor. The multi-camera output resolution configuration includes the output image size of each camera in the current multi-camera combination. The height and width of each output image size are integer multiples of the VAE spatial downsampling factor, and the sum of the number of pixels corresponding to each output image size does not exceed the total pixel budget. The spatial transformation unit is used to perform spatial transformation on each of the video frames according to the multi-camera output resolution configuration, generate the corresponding VAE input image, and record the spatial transformation metadata corresponding to the spatial transformation. The VAE encoding unit is used to perform VAE encoding on each of the VAE input images, generate latent features, and determine the actual latent space grid corresponding to each of the latent features; The reference grid mapping unit is used to perform reference grid mapping on the potential spatial locations in each of the actual potential spatial grids based on the unified reference grid, each of the actual potential spatial grids and the spatial transformation metadata, and to determine the continuous reference grid coordinates corresponding to each of the potential spatial locations. The position encoding unit is used to generate a spatial rotation position code corresponding to each potential spatial position based on the continuous coordinates of the reference grid corresponding to each potential spatial position, and to use the spatial rotation position code to perform position encoding on the features at the corresponding potential spatial positions in each potential feature.
15. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 13.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 13.
17. A computer program product, characterized in that, Includes a computer program / instruction, which, when executed by a processor, implements the method of any one of claims 1 to 13.