Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2722results about "Image memory management" patented technology

Method for automatically drawing OpenGL program by using Vulkan

The invention discloses a method for automatically drawing an OpenGL (Open Graphics Library) program by using Vulkan. The method comprises the following steps of: creating a context used by the Vulkan, initializing each module, processing an OpenGL instruction related to texture and data buffering, and managing storage of texture and data buffering resources in a video memory; a shader program used by the OpenGL is preprocessed into a format acceptable to Vulkan, and an OpenGL shader program instruction is created and destroyed; processing an OpenGL (Open Graphics Library) instruction related to frame buffering to generate structural body information required by Vulkan dynamic rendering; an OpenGL instruction of the sampler is also created; processing an OpenGL (Open Graphics Library) instruction for creating a vertex input format and managing a vertex data buffer area, and maintaining vertex input information, a vertex buffer area and an index buffer area required by Vulkan; and finally, drawing or calculating, distributing and calling Vulkan on the basis of all the instructions.
Owner:ZHEJIANG UNIV +1

Direct3D memory model compatible method based on adaptive occupied resources

The invention discloses a Direct3D memory model compatible method based on adaptive placeholder resources, which comprises the following steps: establishing resource metadata, accessing a scene library, checking an instruction template and double sandboxes when a DXVK is started, describing related resources by the DXVK through the metadata after a D3D application is started, completing resource mapping and metadata dynamic updating, and allocating the resources to the corresponding sandboxes; the DXVK compiles an application shader code, identifies a resource access instruction, matches an access scene and a check template by combining a pipeline stage and binding slot query metadata, instantiates the check instruction and adds the check instruction to the front of a target instruction; according to the unbound resources, adaptive generation of corresponding occupied resources is carried out, the authority and the state are configured, the unbound resources are bound to a Vulkan descriptor set to replace VKNULLHANDLE, and adaptation of the shader interface and the pipeline state parameters is completed to create PSO (Particle Swarm Optimization); and intercepting an application rendering instruction analysis parameter to construct a command buffer area, submitting the command buffer area to a Vulkan queue to trigger a GPU (Graphic Processing Unit) to execute rendering, and finally completing complete conversion and adaptation from the D3D rendering logic to the Vulkan.
Owner:北京麟卓信息科技有限公司

Lossless and lossy automatic hardware compression in graphics-to-graphics network links

An apparatus to facilitate lossless and lossy automatic hardware compression in graphics-to-graphics network links is disclosed. The apparatus includes compressor / decompressor circuitry (CDC) integrated with physical layer (PHY) intellectual property (IP) hardware circuitry for a graphics processor unit (GPU)-to-GPU communication link communicably coupling a first GPU to one or more other GPUs, the CDC to: receive a data message from the first GPU, wherein the data message is in an uncompressed format; determine that a compression process is to be applied to the data message; apply the compression process to the data message to generate a compressed data message; and cause a GPU link IP hardware circuitry that comprises the PHY IP hardware circuitry to transmit the compressed data message over the GPU-to-GPU communication link.
Owner:INTEL CORP

Data processor, data processing method, electronic device and storage medium

Provided in the present disclosure are a data processor, a data processing method, an electronic device and a non-transitory computer-readable storage medium. The data processor comprises a tensor operation unit and N compute units, wherein the tensor operation unit is configured to execute tensor computation on input data, so as to obtain tensor computation results; and the N compute units are configured to execute at least one of a vector operation of the tensor computation results and the generation of input data, wherein first data transmission channels are provided between the tensor operation unit and at least some compute units among the N compute units, and the first data transmission channels are used for directly providing the tensor computation results to the compute units, and directly providing the input data from the compute units to the tensor operation unit. By means of the data processor, access to a global memory can be reduced, the waste of resources is reduced, the data transmission delay is shortened, the overall efficiency of operators is significantly improved, the strong computing power of a tensor operation unit itself is effectively used, and the computation efficiency of the data processor is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Tensor Memory Accelerator Enhancements

One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster including a plurality of graphics cores and tensor processing circuitry. The tensor processing circuitry includes a local memory, a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply and accumulate operation, and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory coupled to the memory interface and the local memory. The tensor data movement accelerator includes circuitry configured to translate the tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Rendering equipment based on three-dimensional Gaussian sputtering

The invention discloses rendering equipment based on three-dimensional Gaussian sputtering. The hardware architecture includes a memory interface, a pre-processing engine, a cardinality ordering engine, a rasterization engine, and the like. The preprocessing engine converts the three-dimensional Gaussian point into a two-dimensional representation and generates a key value pair containing a tile identifier and depth information; the cardinal number sorting engine performs staged efficient sorting through a parallel processing mechanism and a BRAM alternate working mode; and the rasterization engine performs rendering processing by adopting a full-pipeline design and an early-stage stopping mechanism. According to the method, parallel execution of depth sequencing and rasterization is achieved, memory access and calculation redundancy are reduced through optimized data flow design, an efficient and low-power-consumption hardware acceleration solution is provided for three-dimensional Gaussian sputtering rendering, and the method is particularly suitable for application scenes such as virtual reality, games and scientific visualization requiring real-time rendering.
Owner:TSINGHUA UNIVERSITY

Apparatus and method for block-friendly ray traversal

Apparatus and method for efficient storage of BVH nodes in blocks. For example, one embodiment of an apparatus comprises: bounding volume hierarchy (BVH) construction circuitry to construct a BVH based on primitives of a graphics scene; and block allocation hardware logic coupled to or integral to the BVH construction circuitry, the block allocation hardware logic to allocate a plurality of nodes of the BVH into a plurality of blocks for storage in a cache or memory subsystem, the block allocation hardware logic to maximize a number of blocks which include a leading parent node and one or more corresponding child nodes of the plurality of nodes.
Owner:INTEL CORP

Digital twin platform GPU rendering resource dynamic scheduling method and system

The invention relates to a digital twin platform GPU rendering resource dynamic scheduling method and system, and the method comprises the steps: S1, collecting a rendering task feature signal of a current frame in real time, and collecting a GPU multi-dimensional resource consumption signal; s2, based on the rendering task feature signal, the historical resource consumption signal and the scene dynamic change signal, generating a GPU resource demand prediction signal and a load fluctuation trend signal of a future frame; s3, predicting a signal, a task dependency relationship signal, a user priority signal and a real-time resource bottleneck type signal according to the GPU resource demand; s4, executing the dynamic resource allocation instruction signal; and S5, generating a parameter self-optimization signal for updating the prediction and scheduling logic of a subsequent frame according to the actual resource consumption signal of the current frame and the scheduling effect evaluation signal. According to the dynamic scheduling method and system for the GPU rendering resources of the digital twin platform, the problem that the GPU resource utilization rate is low and the real-time performance is difficult to consider at the same time under the dynamic load can be solved.
Owner:ZHONGKE HUIZHI (BEIJING) TECH CO LTD

High-resolution machine vision detection method and system for precise chromatic aberration detection and medium

The invention provides a high-resolution machine vision detection method and system for precise chromatic aberration detection and a medium, and the method comprises the steps: carrying out the calibration of a camera based on a camera calibration tool, and synchronously calibrating a light source parameter and a color parameter; setting shooting parameters based on the calibrated camera, obtaining a workpiece image in real time, preprocessing the workpiece image, mapping the preprocessed image to a standard color space from an original RGB space, and extracting color features of the preprocessed image; comparing the color feature with the color feature of the standard sample, calculating a color difference index, and comparing with the color difference index based on a set color difference threshold to obtain a detection result; by calibrating various parameters of the camera, the visual detection precision is ensured, and by performing color space mapping on the workpiece image and analyzing the difference between the color difference index of the color feature and the color difference threshold value, the color abnormal area is accurately analyzed, and the detection precision of the abnormal color is improved.
Owner:HANGZHOU HUICUI INTELLIGENT TECH CO LTD

Multi-cluster heterogeneous computing power scheduling method for AI training reasoning task

The invention discloses a multi-cluster heterogeneous computing power scheduling method oriented to an AI training reasoning task, and belongs to the technical field of computers. According to a scheduling mechanism with issuing of the AI task as a core, the state of a resource pool is monitored in real time, resources are dynamically allocated, it is ensured that the AI task is efficiently executed, the use efficiency of the whole resource pool is improved, and the load state of an AI application example is detected in real time; through an elastic telescoping function, AI application examples are automatically expanded and shrunk, stable operation of AI tasks is ensured, resources of a plurality of clusters are integrated into a unified resource pool, centralized management and cooperative scheduling of the resources are achieved, the load condition of each cluster is monitored in real time, load balancing is automatically carried out among the clusters, and the AI tasks are reasonably distributed to different clusters. And the GPU resources of the specified model are accurately allocated to the AI task according to the actual demand of the AI task, so that the situation that the AI task application runs on non-optimal resources is avoided, and the efficient running of the AI task is ensured.
Owner:SICHUAN HUIXIN INTELLIGENT COMPUTING TECHNOLOGY CO LTD

Memory writing method and device, storage medium and program product

The invention discloses a memory writing method and device, a storage medium and a program product, and relates to the technical field of memory access, and the method comprises the following steps: determining a memory step length of an initial image tensor corresponding to a target image, determining a target vectorization width according to the memory step length and a memory data width of a target graphics processor, and determining a target data dimension of the target graphics processor, determining a target thread grid corresponding to the initial image tensor according to the target data dimension, performing image tensor slicing by using the target thread grid to obtain a target image tensor, and writing the target image tensor into a memory. According to the method, the memory layout (step length) of the input tensor can be dynamically analyzed, so that the optimal vectorization width is adaptively determined, the memory access efficiency is maximized, slicing is performed according to the data dimension when the target graphics processor outputs the data, the method has adaptability and high efficiency, the GPU video memory bandwidth is fully utilized, and the memory bandwidth utilization rate is improved.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Self-adaptive segmentation method and system for lesion area of seminal vesicle endoscope image

The invention discloses a lesion area self-adaptive segmentation method and system for a seminal vesicle endoscope image, particularly relates to the field of medical image processing, is used for solving the problems of geometric distortion and artifacts in the seminal vesicle endoscope image, and aims to eliminate geometric deviation caused by thick layer sampling through synchronous acquisition and attitude correction. Then, a resampling strategy is adjusted in a self-adaptive mode through key geometric features, the problems of inter-layer artifacts and resolution imbalance are effectively weakened, a lesion segmentation network is optimized through smooth regularization and geometric constraint, continuity and geometric accuracy of lesion boundaries are ensured, finally, the accurate lesion mask is dynamically overlaid to a real-time frame stream, and the real-time frame stream is obtained. A quantitative basis is provided for biopsy path planning and photodynamic dose scheduling; smooth and continuous images are completed and output in a strict time window, the perception ability of an operator to tiny pathological changes is enhanced, meanwhile, the method is suitable for various endoscope devices, motion blur and light spot artifacts are restrained, and focus details are kept clear.
Owner:SECOND AFFILIATED HOSPITAL OF COLLEGE OF MEDICINEOF XIAN JIAOTONG UNIV

Vertical direction pixel parallel depth operation implementation method and device, medium, equipment and product

The invention discloses a depth operation implementation method and device for parallel pixels in the vertical direction, a medium, equipment and a product. The method comprises the following steps: mapping each channel group with a thread bundle containing NT threads; acquiring jumping step length of vertical direction access and parameter configuration of a sliding window operator; performing register resource allocation for each thread to form a sliding window data cache region; determining NH row starting point coordinates of the current scanning round; when the current scanning round is executed, in each column slice participating in calculation, pixel groups at NT row and column positions are scanned from each row starting point coordinate with the jump step length as the span, and the pixel groups are loaded to register groups of all threads in corresponding thread bundles in sequence, so that each thread can process the calculation task of the corresponding local sliding window view. According to the method, thread-level independent calculation task allocation is realized through a pixel parallel strategy among threads, and allocated hardware resources can be utilized to the maximum extent.
Owner:SHANGHAI BIREN TECH CO LTD

Multi-core GPU cache division method, architecture and data access method

The invention provides a multi-core GPU cache division method, architecture and data access method, relates to the technical field of graphics processor design, and aims to realize cross-core access performance system-level optimization and improve the performance of processing large-scale image data by introducing a cache allocation mechanism based on topology awareness. The method comprises the steps that a multi-core-grain GPU topological structure comprising a target computing core grain and at least one associated computing core grain is obtained, a to-be-divided far-end data cache of the target computing core grain is used for caching image data from a far-end video memory, and the to-be-divided far-end data cache is connected with the associated computing core grain through an inter-core-grain interconnection network; if the number of the calculation core particles contained in the multi-core-particle GPU does not exceed a first preset threshold value, the inter-core-particle hop count between each associated calculation core particle and a target calculation core particle is determined based on the topological structure, and then the corresponding capacity of the associated calculation core particles in the to-be-divided cache of the target calculation core particle is determined; and dividing the to-be-divided cache according to the capacity corresponding to each associated calculation core particle.
Owner:BEIHANG UNIV

Combined MX and sparsity representation

One embodiment provides a graphics processor comprising a memory interface and a processing cluster array including a plurality of processing resources interconnected via a switched interconnect network, at least one of the plurality of processing resources including a matrix accelerator configured to execute an instruction to perform a multi-dimensional sparse matrix multiply and accumulate operation having input in a sparse microscaling format including merged sparsity and scaling metadata.
Owner:INTEL CORP

Intermediate noise retrieval for image generation

A method, apparatus, non-transitory computer readable medium, apparatus, and system for image processing include obtaining an input prompt and retrieving an intermediate noise state based on a similarity between the input prompt and a candidate prompt corresponding to the intermediate noise state. An image generation model generates a synthetic image based on the input prompt and the intermediate noise state.
Owner:ADOBE INC

Frame rate improving method, electronic equipment and storage medium

The embodiment of the invention discloses a frame rate improving method, electronic equipment and a storage medium, and relates to the technical field of display processing. The method comprises the steps that when it is determined that the frame rate of a game picture needs to be increased, a frame insertion rendering thread is created; generating a frame insertion image in a frame insertion rendering thread according to the frame rendering data; and inserting the frame insertion image before or after an original image generated by the main rendering thread according to the frame rendering data. Therefore, when the frame rate of the game picture needs to be improved, the electronic equipment can create a frame insertion rendering thread, and the frame insertion rendering thread can share part of rendering tasks of the main rendering thread, so that the main rendering thread can have more time to process complex scenes, perform real-time interaction and the like; according to the technical scheme, jamming, delay and the like caused by overload of the main rendering thread are reduced, so that rendering of the game picture can maintain a stable target frame rate, the rendering effect of the game picture is improved, and then the game experience feeling of a user is enhanced.
Owner:HONOR DEVICE CO LTD

Data processing method and device, processor and electronic equipment

The invention relates to the technical field of artificial intelligence chips, and discloses a data processing method and device, a processor and electronic equipment. Determining a first jump step length required for rearranging the block data based on the block size of the input data in the first dimension; secondly, performing jump reading on the block data according to a first jump step length, rearranging and loading read input elements into a thread private register so as to realize transposition storage of the block data in the thread private register, and mapping a thread direction on a second dimension so as to obtain a data basis of convolution correlation calculation; convolution correlation calculation is executed based on the transposed data in the thread private register and the weight data in the thread private register, repeated data reading in the calculation process is reduced during sliding window calculation, and an operator is converted into calculation performance bottleneck operation from memory access performance bottleneck operation.
Owner:SHANGHAI BIREN TECH CO LTD

Hardware compression of sparse matrix content

The disclosure describes hardware compression of sparse matrix content. One embodiment provides a graphics processor, the graphics processor comprising: a base die, the base die comprising a plurality of chiplet slots; and a plurality of chiplets, the plurality of chiplets being coupled with the plurality of chiplet slots. At least one chiplet of the plurality of chiplets comprises: a graphical core cluster comprising a plurality of processing elements; a shared local memory coupled with the plurality of processing elements; a plurality of matrix engines coupled to the shared local memory; and codec circuitry coupled with the shared local memory and the plurality of matrix engines. Codec circuitry is configured to decode matrix data stored in a first format in a shared local memory into a second format for consumption by a plurality of matrix engines.
Owner:INTEL CORP

Data loading method, data storage method, processor, electronic equipment and medium

The invention provides a data loading method, a data storage method, a processor, electronic equipment and a medium. The data loading method is used for loading a to-be-processed tensor from an original tensor of a memory to a cache region, and comprises the following steps: determining a plurality of requests for loading the to-be-processed tensor in combination with the number of pixels included in the to-be-processed tensor, a data storage format of the original tensor and an initial coordinate of the to-be-processed tensor in a coordinate system determined by the original tensor, the plurality of requests are used for sequentially acquiring data of the original tensor from the original tensor by taking the starting coordinate as a starting point until the number of the acquired data is equal to the number of pixels included in the tensor to be processed, and the data loaded by each request belong to the data range of the original tensor or do not belong to the data range of the original tensor.
Owner:SHANGHAI BIREN TECH CO LTD

Data loading method, data storage method, processor, electronic equipment and medium

The invention provides a data loading method, a data storage method, a processor, electronic equipment and a medium. The data loading method comprises the steps of determining a plurality of first requests for loading a to-be-processed tensor based on an obtained loading mode of the to-be-processed tensor, a shape and a size of the to-be-processed tensor, a data storage format of the to-be-processed tensor and an initial coordinate of the to-be-processed tensor in a coordinate system determined by an original tensor; sequentially sending a first request in the kth loading request group in the X loading request groups, and sequentially obtaining data from the storage space to obtain first data returned by the kth loading request group; in response to the situation that the size of the first data returned by the kth loading request group is not aligned according to the multiple of Y, the first data is filled to obtain second data, and the size of the second data obtained through filling is aligned according to the multiple of Y; and writing the second data into the cache region. According to the data loading method, the complexity of instructions is reduced, and the overall performance of the processor is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Tensor data moving accelerator

The invention relates to a tensor data movement accelerator. One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster comprising a plurality of graphics cores and tensor processing circuitry. The tensor processing circuit includes: a local memory; a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply-accumulate operation; and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory and a local memory coupled to the memory interface. The tensor data moving accelerator includes circuitry configured to convert tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Nerve radiation field rendering method based on dynamic hash coding

The invention discloses a neural radiation field rendering method based on dynamic hash coding, and the method comprises the steps: employing the feature sequence data as the input, calculating the density value and color value of each sampling point through the forward propagation of a neural network, carrying out the volume rendering integral operation according to the ray tracing principle in the direction of a ray, and obtaining the feature sequence data; judging a final color output result of the current pixel point; according to an error value between the color output result and a real image, updating a network parameter weight through a back propagation algorithm, and if the error value is greater than a convergence threshold, continuing to iterate the training process to adjust a feature coding strategy to obtain an optimized neural radiation field model parameter; and after the rendering performance configuration parameters are obtained, optimizing a storage allocation strategy of feature data through a memory pool management mechanism, and if the current memory occupancy rate exceeds a safety threshold, starting a data compression algorithm to reduce the storage space requirement, and obtaining a real-time rendering output result. According to the invention, high-quality real-time rendering of the dynamic scene is realized.
Owner:ZHEJIANG UNIV OF TECH

GPUBox hardware decoupling system based on Retimer card and PCIeSwitch chip

The invention discloses a GPU Box hardware decoupling system based on a Retimer card and a PCIe Switch chip, and belongs to the technical field of computer hardware architecture and high-speed interconnection. According to the system, a Retimer card and a PCIe Switch chip are integrated in an independent GPU Box, and a decoupling link of a CPU server and a GPU acceleration card is constructed; the Retimer card realizes 30-meter long-distance PCIe signal transmission and breaks through physical distance limitation; the PCIe Switch chip pools GPU resources through a dynamic routing and MRIOV technology, supports flexible allocation of computing power by multiple servers, and realizes Peer-to-Peer direct connection communication between GPUs. Aiming at a large model reasoning scene, the system optimizes KV cache bandwidth allocation and video memory and memory cooperative scheduling, so that the 100B parameter model reasoning throughput is greatly improved; and meanwhile, the usability of the system is greatly improved through fault isolation and hot plug design. According to the method, the problems of physical binding of the CPU and the GPU, limited transmission distance, rigid resource allocation and the like in a traditional architecture are solved, and the method is suitable for large-scale AI calculation and distributed GPU cluster deployment.
Owner:HEFEI FENGZHIYI SEMICON CO LTD

Speculative execution of kernel programs in a chiplet based architecture

One embodiment provides a multi-chiplet graphics processor comprising a plurality of chiplets, where a chiplet of the plurality of chiplets comprise a memory interface, processing resources configured to execute threads of a kernel, and thread dispatch circuitry to facilitate dispatch of threads of the kernel to the processing resources. The processing resources are configured to execute threads of a first kernel, receive dispatch of threads of a second kernel for execution before completion of the first kernel as threads of the first kernel retire, execute a first phase of the second kernel during completion of execution of the first kernel, via a thread of the first kernel, signal an event via an uncached write to a global memory, and execute a second phase of the second kernel based on detection of the event via an uncached read from the global memory.
Owner:INTEL CORP

Data loading method, data storage method, processor, electronic equipment and medium

The invention discloses a data loading method, a data storage method, a processor, electronic equipment and a medium. The data loading method comprises the steps that a first coordinate value and a first size are obtained, target data comprises a plurality of data parts, the coordinate ranges of all the data parts on the channel number dimension in a coordinate system determined by an original tensor are the same, and all the data parts are from the first coordinate value to a first channel coordinate; the difference between the first coordinate value and the first channel coordinate plus 1 is equal to a first size; a coordinate information table is obtained, the coordinate information table is used for indicating initial coordinates of each data part on at least one dimension in a coordinate system, and the at least one dimension is other dimensions except the channel number dimension in multiple dimensions included in the coordinate system; based on the first coordinate value, the first size and the coordinate information table, determining a plurality of requests for loading the target data; and sending a plurality of requests in sequence, and writing the sub-data returned by each request into the cache region in sequence to load the target data to the cache region.
Owner:SHANGHAI BIREN TECH CO LTD

Bindless thread dispatch mid-thread preemption on a graphics processor

Techniques are provided to enable mid-thread (instruction level) preemption in a graphics processor without requiring software intervention by a graphics driver associated with the graphics processor. Hardware based mid-thread preemption is facilitated using the thread dispatch hardware of the graphics processor to trigger execution of a kernel program by the graphics processor that saves the thread state of preempted processing resources and facilitates the subsequent restoration of that thread state to the same or a different set of processing resources.
Owner:INTEL CORP

Multi-view collaborative 3D Gaussian splash optimization method and system

The invention discloses a multi-view collaborative 3D Gaussian splash optimization method and system, and the method comprises the steps: constructing a multi-level heterogeneous video memory pool, and dynamically dividing a video memory in 3D Gaussian reconstruction into a view exclusive memory block and a global shared memory pool; a mixed rendering-gradient pipeline is designed, and hardware-level pipeline parallelism in 3D Gaussian reconstruction is realized through a double-buffer asynchronous switching mechanism based on a CUDA Warp-level parallel primitive fusion forward rendering and back propagation thread group; performing multi-view gradient joint optimization, screening an effective gradient path in 3D Gaussian reconstruction based on the visibility mask matrix, and performing projection error weighted fusion on a multi-view gradient tensor; and implementing a multi-modal densification decision, generating a 3D Gaussian candidate splitting position in 3D Gaussian reconstruction through Monte Carlo sampling, calculating a joint optimization objective function by combining a multi-view projection residual error and a gradient contribution factor, and finally realizing 3D Gaussian reconstruction. According to the invention, high-precision and low-delay large-scale scene real-time rendering and training can be realized.
Owner:ZHEJIANG UNIV

NPU-based SiamFC-pytorch tracking model pre-processing and post-processing acceleration method and system

The invention aims to provide a SiamFC-pytorch tracking model pre-processing and post-processing acceleration method and a SiamFC-pytorch tracking model pre-processing and post-processing acceleration system based on an NPU (Network Processing Unit). The system is realized based on a Yulong810SOC, and comprises a video input module (1), a hardware decoding module (2), a detection module (3), a tracking module (4) and a turntable control module (5), the method comprises the following steps: a, a video input module receives a video stream, and the video stream is decoded into an original image in a YUV format through a hardware decoding module; b, the detection module runs a YOLOv5 model through a neural network engine of the NPU; c, an NPU pre-processing and post-processing unit of the tracking module executes template branch pre-processing, search branch pre-processing and search branch post-processing on the YUV image and the target initial bounding box; and d, the turntable control module adjusts the posture of the camera turntable according to the real-time coordinates of the target. The method is applied to the technical field of computer vision and artificial intelligence hardware acceleration.
Owner:ZHUHAI ORBITA AEROSPACE SCI TECH CO LTD