ISP (Internet Service Provider) and NPU (Network Processing Unit)-based intelligent visual chip fusion architecture and construction method

By integrating the ISP and NPU into a smart vision chip architecture, and employing direct data channels and dynamic resource scheduling, the problems of low data interaction efficiency and low hardware resource utilization in traditional architectures are solved. This enables efficient visual data processing and analysis, adapts to complex environments, and improves system performance and security.

CN121545015APending Publication Date: 2026-02-17STATE GRID HENAN INFORMATION & TELECOMM CO +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511770142.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Traditional ISP and NPU architectures suffer from low efficiency in inter-frame data interaction, insufficient environmental adaptation and intelligent analysis collaboration, and low hardware resource utilization. This results in excessive data transmission volume, low processing efficiency, and wasted hardware resources, making it difficult to meet the imaging quality and analysis accuracy requirements in complex environments.

Method used

The system adopts an intelligent vision chip fusion architecture that integrates ISP and NPU. It enables bidirectional data interaction through a direct data channel. The intra-frame collaborative controller dynamically delineates the region of interest, the inter-frame optimization module identifies dynamically changing regions, and the imaging parameters of the ISP are adjusted in conjunction with the inter-frame trend prediction signal of the NPU. The dynamic resource scheduling unit monitors the load status in real time and migrates resources to achieve data transmission optimization and improved resource utilization.

Benefits of technology

It significantly reduces data transmission latency and throughput, improves imaging quality and analysis accuracy, enhances hardware resource utilization, reduces energy consumption, adapts to processing needs in different scenarios, and strengthens data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545015A_ABST
    Figure CN121545015A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent visual chips, in particular to an intelligent visual chip fusion architecture based on ISP and NPU and a construction method. The architecture comprises an ISP (Internet Service Provider), an NPU (Network Processing Unit), an intra-frame cooperative controller, an inter-frame optimization module and a dynamic resource scheduling unit. Bidirectional data interaction is carried out between the ISP and the NPU through a direct connection data channel; the intra-frame cooperative controller dynamically defines a self-adaptive region of interest (ROI) according to feature position information fed back by the NPU, and only transmits preprocessed feature data of the ROI to the NPU; the inter-frame optimization module analyzes, identifies and only transmits visual feature retention data of a dynamic change area according to continuous inter-frame difference images, and meanwhile, adjusts a next frame imaging parameter of the ISP; and the dynamic resource scheduling unit dynamically migrates the computing resources or the storage bandwidth when a preset condition is met. According to the invention, deep collaboration of the ISP and the NPU is realized, and the processing efficiency and performance of the visual chip are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent vision chip technology, specifically to an intelligent vision chip fusion architecture and construction method based on ISP and NPU. Background Technology

[0002] With the rapid development of artificial intelligence technology, intelligent vision systems are increasingly being used in industrial inspection, security monitoring, autonomous driving, and other fields. As the core of an intelligent vision system, the processing power and efficiency of the vision chip directly affect the overall performance of the system. Traditional vision chip architectures typically design the image signal processor (ISP) and neural network processor (NPU) as independent modules, with the two interacting through shared memory.

[0003] Currently, various vision chip architecture solutions are available on the market. For example, CN102665049B discloses a vision image processing system based on a programmable vision chip. This system includes an image sensor and multi-level parallel digital processing circuits, enabling high-speed, high-quality image acquisition and multi-level parallel image processing. It allows for the implementation of various high-speed intelligent vision applications through programming. CN117689577A proposes a high-definition camera video processing method and system. This method performs video frame extraction, edge enhancement processing, image matrix segmentation, visual channel decomposition, and independent neural network encoding on high-definition camera videos, achieving efficient and clear video processing.

[0004] In the area of ​​image processing in complex environments, CN120876278A discloses an adaptive contrast-enhanced video fusion algorithm for complex logistics scenarios. This algorithm extracts environmental features from an initial video dataset, uses a convolutional neural network to analyze inter-frame illumination changes and object movement speeds to determine the current scene environment, and performs detail preservation enhancement on foreground objects. CN120219330A proposes a visual image monitoring method and electronic chip. It acquires raw image streams using a biomimetic retina-based complementary metal-oxide-semiconductor sensor, generates dynamic sensing event streams, and performs target detection and trajectory prediction processing.

[0005] In terms of chip resource optimization, CN120510824A discloses a power consumption optimization control system for high refresh rate driver chips. This system models the refresh requirements of image regions by dividing the region and modeling the graph structure, combined with attention mechanism and loss feedback, to achieve differentiated refresh, effectively avoid invalid refresh in low-change areas, improve the content awareness capability of the refresh strategy, and thus reduce overall power consumption.

[0006] However, the following problems still exist in the existing technology: First, the data interaction between the ISP and NPU in the traditional architecture is inefficient. After the ISP completes the processing of a single frame of image, it needs to store the complete image data in shared memory, and then the NPU reads the data from the shared memory for analysis. This method has redundant storage and retrieval steps. At the same time, the existing solution does not optimize the data transmission content for the inter-frame correlation of visual data, resulting in excessive data transmission volume, which affects the real-time performance of intelligent analysis.

[0007] Second, the existing architecture lacks sufficient collaboration between the ISP and NPU. The ISP's image optimization does not take into account the NPU's analysis needs, and the NPU's analysis results are not fed back to the ISP to guide subsequent optimization. This fragmented approach makes it difficult for the system to simultaneously achieve both image quality and analysis accuracy in complex environments, especially in complex scenes such as low light, backlight, rain, and fog.

[0008] Third, the existing vision chip architecture has low hardware resource utilization. Traditional architectures do not dynamically allocate hardware resources according to the task load of the ISP and NPU. When one processor needs a lot of computing resources, the idle resources of the other cannot be allocated and used, resulting in a waste of overall chip hardware resources, increased device power consumption, and is not conducive to the long-term operation of mobile or edge computing devices.

[0009] Furthermore, existing technologies lack adaptive processing mechanisms for different scenarios, making it impossible to dynamically adjust processing strategies based on environmental changes and task requirements, thus making it difficult to maintain analytical accuracy while ensuring processing efficiency. At the same time, the data security mechanisms in existing architectures are relatively simple and cannot meet the ever-increasing demands for data security. Summary of the Invention

[0010] To address the technical problems of low inter-frame data interaction efficiency, insufficient environmental adaptation and intelligent analysis collaboration, and low hardware resource utilization in traditional ISP and NPU architectures, and to achieve improved data interaction efficiency, optimized imaging and analysis collaboration, optimized hardware resource utilization, and enhanced scene adaptability and security, this invention provides an intelligent vision chip fusion architecture and construction method based on ISP and NPU.

[0011] According to one aspect of the present invention, a smart vision chip fusion architecture based on ISP and NPU is provided, comprising: ISP, NPU, intra-frame co-controller, inter-frame optimization module, and dynamic resource scheduling unit; the ISP and NPU perform bidirectional data interaction through a direct data channel rather than shared memory; the intra-frame co-controller is used to dynamically delineate adaptive regions of interest (ROIs) in the image based on feature location information fed back by the NPU, and simultaneously transmit only the preprocessed feature data of the adaptive ROIs to the NPU through the direct data channel; the inter-frame optimization module is used to identify and transmit only the visual feature retention data of dynamically changing regions based on differential image analysis between consecutive frames, and actively adjust the imaging parameters of the ISP for the next frame in combination with the inter-frame trend prediction signal output by the NPU; the dynamic resource scheduling unit is used to monitor the load status of the ISP and NPU in real time, and dynamically migrate computing resources or storage bandwidth between the ISP and the NPU when a preset trigger condition is met.

[0012] In some optional implementations of certain embodiments, the rules for defining the adaptive region of interest (ROI) include: defining the ROI region with the coordinates of the fault point fed back by the NPU as the center and an adaptive radius, wherein the adaptive radius is dynamically adjusted according to the device size.

[0013] In some optional implementations of certain embodiments, the intra-frame cooperative controller has a built-in environment task adaptation rule mapping table, which includes: when a night vision scene is detected, the ISP enables 3D noise reduction and gain enhancement, and the NPU enables the YOLOv5s-DeepSort model and sets the IoU threshold; when a backlight scene is detected, the ISP enables local exposure equalization and dynamic range expansion, and the NPU enables anti-halo feature enhancement recognition mode; when a rain or fog scene is detected, the ISP enables a transmittance estimation defogging algorithm, and the NPU enables a fog and haze robust feature extraction network.

[0014] In some alternative implementations of certain embodiments, the differential image analysis calculates differential images of consecutive frames and binarizes the differential images using an adaptive double thresholding method to separate dynamic regions.

[0015] In some optional implementations of some embodiments, the inter-frame optimization module adopts a visual feature preservation compression strategy based on structural similarity, specifically: for pixel regions in the neighborhood of the fault point where the gradient change rate exceeds a preset threshold, the original precision is forcibly preserved, while lossy compression transmission is used for the remaining regions.

[0016] In some optional implementations of certain embodiments, the inter-frame trend prediction signal is the continuous inter-frame target state change trend output by the NPU. The ISP uses a linear prediction model based on this trend to adjust the exposure time, gain, or white balance parameters in advance to suppress brightness or contrast fluctuations in the next frame image.

[0017] In some optional implementations of certain embodiments, the preset triggering conditions of the dynamic resource scheduling unit include: when the ISP computing unit occupancy rate is continuously higher than a first threshold and the NPU task queue is empty, triggering the temporary migration of some NPU computing units to the ISP; when the NPU task queue length exceeds a second threshold and the ISP storage bandwidth idle rate is higher than a third threshold, triggering the reuse of ISP storage bandwidth to accelerate NPU data reading.

[0018] In some optional implementations of certain embodiments, an integrated secure data bus is also included, used to connect the ISP, NPU, intra-frame cooperative controller, inter-frame optimization module and dynamic resource scheduling unit to realize data transmission.

[0019] In some optional implementations of certain embodiments, the system further includes an encoding / decoding module and a security control unit. The integrated data bus directly transmits the image processed by the ISP to the encoding / decoding module, and the NPU's analysis results are encrypted and output uniformly by the security control unit.

[0020] According to a second aspect of the present invention, a method for constructing a smart vision chip fusion architecture based on an ISP and an NPU is provided, comprising: setting up a direct data channel between the ISP and the NPU, wherein the ISP and the NPU perform bidirectional data interaction through the direct data channel rather than shared memory; dynamically delineating an adaptive region of interest (ROI) in an image based on feature location information fed back by the NPU through an intra-frame collaborative controller, while transmitting only the preprocessed feature data of the ROI to the NPU through the direct data channel; identifying and transmitting only the visual feature retention data of dynamically changing regions based on differential image analysis between consecutive frames through an inter-frame optimization module, and actively adjusting the imaging parameters of the ISP for the next frame in combination with the inter-frame trend prediction signal output by the NPU; and monitoring the load status of the ISP and the NPU in real time through a dynamic resource scheduling unit, and dynamically migrating computing resources or storage bandwidth between the ISP and the NPU when a preset trigger condition is met.

[0021] The intelligent vision chip fusion architecture and construction method based on ISP and NPU provided in this application have the following beneficial effects: By using intra-frame direct connection channels and inter-frame differential transmission mechanisms, data transmission links and transmission volume are reduced. In 1080P video processing scenarios for power monitoring, the data transmission latency between ISP and NPU is reduced, and the data transmission volume is decreased, solving the problem of NPU lag while waiting for ISP data and ensuring the real-time performance of power equipment status analysis. Through intra-frame analysis result feedback guiding ISP optimization and inter-frame trend prediction optimizing the next frame imaging, deep collaboration between the two is achieved. In night vision scenarios for power monitoring, the accuracy of NPU in identifying equipment fault points is improved. In backlight scenarios, the ISP output image... The clarity of the core area of ​​the device is improved, while the fluctuation of the NPU analysis results is reduced. Through dynamic resource scheduling, the reuse of ISP and NPU resources is realized. In the typical operating scenarios of power monitoring equipment, the overall hardware resource utilization of the chip is improved and the power consumption of the equipment is reduced, which meets the low power consumption and long battery life operation requirements of power monitoring equipment. At the same time, it can reduce the size of chip hardware and reduce manufacturing costs. The environmental task adaptation rules can automatically adjust the working mode of ISP and NPU for typical power monitoring environments such as night vision, backlight, rain and fog. The fusion architecture seamlessly connects with the dedicated security control unit through an integrated data bus to ensure that imaging data and analysis results are encrypted and protected during transmission, which meets the requirements of the power digital vision security system. Attached Figure Description

[0022] Figure 1 This is an architecture diagram of the intelligent vision chip fusion architecture based on ISP and NPU in an embodiment of the present invention.

[0023] Figure 2 This is a flowchart of the intelligent vision chip fusion architecture based on ISP and NPU in an embodiment of the present invention.

[0024] Reference numerals: 1-ISP, 2-Intra-frame Cooperative Controller, 3-NPU, 4-Inter-frame Optimization Module, 5-Dynamic Resource Scheduling Unit, 6-Integrated Secure Data Bus, 7-Security Control Unit, 8-Encoding / Decoding Module. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0026] Example 1 Reference Appendix Figure 1This application provides an intelligent vision chip fusion architecture based on ISP and NPU. The architecture mainly includes an image signal processor ISP1, a neural network processor NPU3, an intra-frame co-controller 2, an inter-frame optimization module 4, a dynamic resource scheduling unit 5, and an integrated secure data bus 6. These components achieve efficient visual data processing and analysis through innovative architecture design.

[0027] In this architecture, a direct data channel is established between ISP1 and NPU3, replacing the traditional shared memory method and enabling bidirectional data interaction. This direct connection significantly reduces data transmission latency and improves processing efficiency. The content transmitted through the direct data channel may include brightness, edge, and motion vector feature maps output by ISP1, rather than complete pixel data, which further reduces transmission bandwidth requirements and improves data processing efficiency. In this embodiment, when ISP1 processes an image, it can divide the image into core and non-core regions. The preprocessed data of the core region is transmitted to NPU3 in real time through the direct channel, while the data of the non-core region is cached locally in ISP1, reducing the amount of data transmission. After receiving the core region data, NPU3 can perform feature analysis in parallel and provide real-time feedback of key feature location information to ISP1, guiding ISP1 to optimize the image quality of that region.

[0028] Based on the feature location information (specifically, the fault point coordinates) fed back by NPU3, the intra-frame cooperative controller 2 dynamically delineates an adaptive Region of Interest (ROI) in the image, centered on that point, and transmits only the pre-processed feature data of the ROI to NPU3 via a direct data connection. The ROI delineation rules include using the fault point coordinates fed back by NPU3 as the center and delineating the ROI region according to an adaptive radius. This adaptive radius is dynamically adjusted according to the device size, ranging from 5×5 to 15×15 pixels. This selective data transmission mechanism significantly reduces the amount of data processing, allowing the system to concentrate computing resources on processing the most critical image regions. Furthermore, the intra-frame cooperative control logic design includes: developing an intra-frame cooperative controller 2 that synchronizes the intra-frame processing timing of ISP1 and NPU3, ensuring that the data transmission of ISP1 matches the rhythm of NPU3 analysis; it also incorporates environmental task adaptation rules. For example, when a night vision scene is detected, the controller automatically triggers the low-noise priority processing mode of ISP1 and the high-sensitivity feature recognition mode of NPU3, ensuring that the processing strategies of both are consistent.

[0029] The intra-frame cooperative controller 2 incorporates an environment task adaptation rule mapping table, which contains configuration rules for ISP1 and NPU3 under various scenarios. When a night vision scene is detected, ISP1 enables 3D noise reduction and gain enhancement, while NPU3 enables the YOLOv5s-DeepSort model and sets the IoU threshold. When a backlight scene is detected, ISP1 enables local exposure equalization and dynamic range expansion, while NPU3 enables an anti-halo feature enhancement recognition mode. When a rain or fog scene is detected, ISP1 enables a transmittance estimation defogging algorithm, while NPU3 enables a fog robust feature extraction network. This scene adaptation mechanism enables the chip to optimize processing strategies for different environmental conditions, improving recognition accuracy.

[0030] The inter-frame optimization module 4 analyzes the differential images between consecutive frames, identifies and transmits only the visual feature retention data of dynamically changing regions, and actively adjusts the imaging parameters of the next frame of ISP1 by combining the inter-frame trend prediction signal output by NPU3. Differential image analysis calculates the differential images of consecutive frames and uses an adaptive double-threshold method to binarize the differential images to separate dynamic regions. This mechanism reduces redundant data processing and improves the system's response speed to dynamic scenes.

[0031] The inter-frame optimization module 4 employs a visual feature-preserving compression strategy based on structural similarity (SSIM). For pixel regions where the gradient change rate in the neighborhood of a fault point exceeds a preset threshold, the original precision is forcibly preserved, while lossy compression is used for transmission in the remaining regions. This differentiated compression strategy reduces the overall data transmission volume while ensuring the image quality of critical areas.

[0032] The inter-frame trend prediction signal output by NPU3 represents the changing trend of the target state between consecutive frames. Based on this trend, ISP1 uses a linear prediction model to adjust exposure time, gain, or white balance parameters in advance to suppress brightness or contrast fluctuations in the next frame. This feedforward adjustment mechanism allows imaging parameters to adapt to scene changes in advance, improving image quality stability.

[0033] The inter-frame correlation data optimization mechanism between ISP1 and NPU3 specifically includes: (1) Inter-frame data compression transmission: Analyze the inter-frame correlation of power monitoring video, extract static and dynamic data in continuous frames, and design an inter-frame differential transmission algorithm. This algorithm first calculates the inter-frame correlation of continuous frames. and difference image The formula is: Subsequently, the adaptive double threshold method is used to binarize the differential image to accurately separate the dynamic region. ISP1 only transmits the updated part of the dynamic data and static data to NPU3 to replace the transmission of the complete frame data. At the same time, the visual feature preservation compression strategy is used to ensure that the key features of the device in the dynamic data are not lost. For example, the subtle changes of the device fault point can be transmitted completely. (2) Visual feature preservation compression strategy: SSIM-based local structural similarity detection is used. For the region where the gradient change rate of the neighboring pixels of the fault point is greater than the threshold, the accuracy is forcibly preserved, and the accuracy of the remaining regions is reduced. (3) Feedback of inter-frame analysis results: After NPU3 completes the analysis of the current frame, it feeds back the inter-frame change trend to ISP1. ISP1 combines the trend to predict the image optimization direction of the next frame. For example, when it is predicted that the light will be enhanced in the next frame, the linear prediction model is used as follows: Here, P represents imaging parameters such as exposure time and gain, and α is a smoothing coefficient. Pre-adjusting exposure parameters avoids brightness fluctuations in consecutive frames, improving video imaging stability. In backlit scenarios of power monitoring, this effectively prevents fluctuations in NPU3 analysis results caused by inconsistent image brightness. This cross-modal temporal prediction feedback closed-loop mechanism uses the inter-frame trend derived from NPU3 analysis as a feedback signal to actively predict and pre-adjust the imaging parameters of the next frame in ISP1. Existing technologies only provide feedback on position or quality gradients, without achieving trend prediction and pre-optimization.

[0034] The dynamic resource scheduling unit 5 monitors the load status of ISP1 and NPU3 in real time, and dynamically migrates computing resources or storage bandwidth between ISP1 and NPU3 when preset trigger conditions are met. The preset trigger conditions include: when the ISP1 computing unit utilization rate is consistently higher than a first threshold and the NPU3 task queue is empty, temporarily migrating some computing units of NPU3 to ISP1 is triggered; when the NPU3 task queue length exceeds a second threshold and the ISP1 storage bandwidth idle rate is higher than a third threshold, reusing the ISP1 storage bandwidth is triggered to accelerate NPU3 data reading. This dynamic resource scheduling mechanism improves the overall resource utilization of the chip and prevents a single processing unit from becoming a system bottleneck.

[0035] Load-aware resource scheduling unit design: An ISP1-NPU3 load monitor was developed to collect real-time task load and hardware resource usage status of both devices; a dynamic resource allocation algorithm was designed, the core of which is a mixed-integer linear programming optimization problem, aiming to minimize the total task completion time while satisfying real-time constraints. Its objective function is: ,in, and These are the estimated completion times for tasks in ISP1 and NPU3, respectively, and their values ​​are affected by the allocated computing resources. and Storage bandwidth impact. The algorithm solves the resource allocation scheme based on real-time load status: when ISP1 is overloaded and NPU3 is idle, some computing resources of NPU3 are temporarily allocated to ISP1 for image optimization; when NPU3 is overloaded, idle storage bandwidth of ISP1 is reused to accelerate data reading, while ensuring that resource allocation does not affect the core tasks of both. The specific triggering conditions of the resource scheduling algorithm include: when the ISP1 computing unit utilization rate is too high and the NPU3 task queue is empty for a period of time, NPU3 computing unit migration is triggered. When the NPU3 task queue is too long and ISP1 storage bandwidth is idle, ISP1 bandwidth reuse is triggered.

[0036] The integrated secure data bus 6 connects ISP1, NPU3, intra-frame co-controller 2, inter-frame optimization module 4, and dynamic resource scheduling unit 5 to achieve data transmission. This bus adopts a secure transmission mechanism to protect the integrity and confidentiality of data during its internal flow within the chip.

[0037] In addition, the intelligent vision chip fusion architecture also includes an encoding / decoding module 8 and a security control unit 7. An integrated data bus directly transmits the image processed by the ISP1 to the encoding / decoding module 8, and the analysis results from the NPU3 are encrypted and output uniformly by the security control unit 7. This design enhances data security and prevents the leakage of sensitive information.

[0038] In this embodiment, the intra-frame coordination module, inter-frame optimization mechanism, resource scheduling unit, and other functional modules of the chip are integrated to design an integrated data bus. Data processed by the intra-frame coordination module can be directly transmitted to the encoding / decoding module 8 for compression. The analysis results of NPU3 can be encrypted and output through the security control unit 7, avoiding security risks during data transmission between multiple modules. Simultaneously, resource sharing among modules is achieved through the bus, further improving the overall operating efficiency of the chip. The encryption mechanism of the integrated data bus is as follows: the NPU3 output results are encrypted using AES-128-GCM, and the key is dynamically derived by the security control unit 7 based on the chip's unique ID and updated every frame.

[0039] Furthermore, this application also involves the construction of a multi-scenario verification environment for architectural function verification and optimization. Specifically, it involves building a test environment simulating power monitoring scenarios, including three typical environments: night vision, backlight, and rain / fog, and inputting power equipment monitoring videos of different resolutions. A three-dimensional verification system is constructed: verification is performed from three dimensions: data interaction efficiency, imaging analysis collaboration effect, and resource utilization. If any dimension fails to meet the standard, optimization is performed by returning to the intra-frame collaborative control logic or inter-frame feedback mechanism until the requirements of the power monitoring scenario are met.

[0040] Through the collaborative work of the aforementioned components, the intelligent vision chip fusion architecture achieves efficient visual data processing, analysis, and transmission, significantly improving system performance and resource utilization while ensuring data security. This architecture is particularly suitable for vision applications with high requirements for real-time performance, accuracy, and security, such as intelligent monitoring, autonomous driving, and industrial quality inspection.

[0041] Example 2 Reference Appendix Figure 2 This embodiment, based on Embodiment 1 above, provides a method for constructing a fusion architecture for an intelligent vision chip based on ISP and NPU. This method effectively improves the overall performance and efficiency of the intelligent vision processing system by optimizing the data transmission mechanism between ISP1 and NPU3, implementing a regionalized processing strategy, and dynamic resource management. The specific implementation process is as follows: S1. A direct data channel is set up between ISP1 and NPU3, and ISP1 and NPU3 perform bidirectional data interaction through the direct data channel instead of shared memory.

[0042] In traditional vision processing architectures, ISP1 and NPU3 typically interact via shared memory, which leads to data transmission latency and bandwidth bottlenecks. This method first establishes a dedicated direct data channel between ISP1 and NPU3. This channel utilizes high-speed serial interface technology, supports bidirectional data transmission, and boasts a bandwidth of up to 12GB / s, significantly exceeding the 3-5GB / s transmission speed of traditional shared memory methods. The direct data channel employs a point-to-point connection architecture, reducing data contention on the system bus and ensuring real-time performance and reliability of data transmission through a dedicated hardware handshake mechanism.

[0043] S2. Based on the feature location information fed back by NPU3, the intra-frame co-controller 2 dynamically delineates the adaptive region of interest (ROI) in the image, and transmits the preprocessed feature data of the ROI to NPU3 only through the direct data channel.

[0044] Intra-frame co-controller 2 receives feature location information from NPU3, which includes the location, size, and importance score of key targets in the current frame image. Based on this information, intra-frame co-controller 2 executes an adaptive Region of Interest (ROI) partitioning algorithm. This algorithm first performs cluster analysis on feature points to identify high-density feature regions; then, it calculates the comprehensive weight of each region based on the importance score of the feature points; finally, it determines the boundary of the final adaptive ROI using a dynamic thresholding method. The shape of the adaptive ROI can be rectangular, polygonal, or irregular, automatically adjusted according to the scene complexity.

[0045] After determining the adaptive Region of Interest (ROI), ISP1 performs high-quality preprocessing only on the image data within the ROI, including noise reduction, color correction, and sharpening, generating preprocessed feature data. This preprocessed feature data is transmitted to NPU3 via a direct data connection, significantly reducing data transmission volume. Actual measurements show a reduction of 65% to 85%, depending on the ROI's proportion.

[0046] S3. Based on the differential image analysis between consecutive frames, the inter-frame optimization module 4 identifies and transmits only the visual feature retention data of dynamically changing areas, and combines the inter-frame trend prediction signal output by NPU3 to actively adjust the imaging parameters of the next frame of ISP1.

[0047] The inter-frame optimization module 4 employs an efficient inter-frame difference algorithm to perform pixel-level comparisons between two consecutive frames, generating a difference image. This difference image undergoes adaptive thresholding to identify regions experiencing significant changes. For these dynamically changing regions, the system extracts and retains only their visual feature data, including edge features, texture features, and color change features, while for static regions, the processing results from the previous frame are reused.

[0048] Simultaneously, the inter-frame optimization module 4 receives the inter-frame trend prediction signal output by the NPU3. This signal contains prediction information such as the target's motion trajectory, velocity changes, and possible new target positions. Based on this prediction information, the inter-frame optimization module 4 actively adjusts the imaging parameters of the ISP1 for the next frame, including exposure time, gain settings, sensitivity, and color balance, to optimize the imaging quality of the predicted area. This feedforward adjustment mechanism enables the system to adapt to scene changes in advance, reducing processing latency and improving image quality.

[0049] S4. The dynamic resource scheduling unit 5 monitors the load status of ISP1 and NPU3 in real time, and dynamically migrates computing resources or storage bandwidth between ISP1 and NPU3 when the preset trigger conditions are met.

[0050] The dynamic resource scheduling unit 5 monitors the load status of ISP1 and NPU3 in real time, including processor utilization, memory usage, processing latency, and queue depth. When the following preset trigger conditions are detected, the system will initiate dynamic resource migration. For example, when the load of ISP1 exceeds 85% and the load of NPU3 is below 40%, some image preprocessing tasks (such as edge detection and simple feature extraction) will be migrated to NPU3 for execution; when the load of NPU3 exceeds 90% and the load of ISP1 is below 50%, some low-complexity neural network inference tasks (such as simple classifiers) will be migrated to the programmable processing unit of ISP1 for execution; when the memory bandwidth utilization of a certain processing unit exceeds 75%, the cache allocation strategy will be dynamically adjusted to allocate some cache resources to high-load units.

[0051] Resource migration employs task decomposition and reorganization techniques, breaking down large processing tasks into multiple subtasks and dynamically allocating them to the most suitable processing units based on the current resource status. During the migration process, the system maintains the continuity of the processing pipeline to ensure uninterrupted data processing.

[0052] Furthermore, the intelligent vision chip fusion architecture of this application achieves system interconnection through an integrated secure data bus 6. The integrated secure data bus 6 connects the ISP1, NPU3, intra-frame collaborative controller 2, inter-frame optimization module 4, and dynamic resource scheduling unit 5, forming a complete intelligent vision processing system. This bus adopts a layered architecture design, including a physical layer, link layer, network layer, and application layer, supporting multiple transmission modes, including point-to-point transmission, broadcast transmission, and multicast transmission.

[0053] The secure data bus integrates a data encryption engine that supports the AES-128-GCM encryption algorithm for real-time encryption of sensitive data. It also implements an access control mechanism, assigning different access permissions to different modules to prevent unauthorized access. The bus protocol supports Quality of Service (QoS) management, dynamically adjusting transmission bandwidth allocation based on data priority to ensure the real-time transmission of critical data.

[0054] In a preferred embodiment, the direct data channel between ISP1 and NPU3 is implemented using a PCIe 4.0 x8 interface, providing a theoretical bandwidth of 16GB / s in both directions and supporting zero-copy data transmission technology to further reduce data transmission latency.

[0055] In another preferred embodiment, the intra-frame cooperative controller 2 employs an attention-based ROI partitioning algorithm. This algorithm calculates the visual saliency and semantic importance of each region of the image to more accurately determine the ROI boundaries, making the ROI regions more closely match the actual target contours, reducing redundant regions, and further improving data transmission efficiency.

[0056] In another preferred embodiment, the inter-frame optimization module 4 combines the scene understanding capabilities of deep learning to establish dedicated parameter adjustment models for different scene types (such as night vision, backlight, rain and fog, etc.), and automatically selects the best parameter adjustment strategy according to scene characteristics to improve the scene adaptability of imaging quality.

[0057] By implementing the above technical solution, the intelligent vision chip fusion architecture constructed using this method reduces processing latency by 45% and energy consumption by 38% compared to traditional architectures, while maintaining or improving visual recognition accuracy. This method is particularly suitable for applications requiring real-time visual processing, such as intelligent security, autonomous driving, and industrial inspection.

[0058] Finally, it should be noted that the above descriptions are merely preferred embodiments of this application, and this application is not limited to the above embodiments. It is understood that other improvements and variations directly derived or conceived by those skilled in the art without departing from the spirit and concept of this application should be considered to be included within the protection scope of this application.

Claims

1. An intelligent vision chip fusion architecture based on ISP and NPU, characterized in that, Comprise: ISP, NPU, intra-frame cooperative controller, inter-frame optimization module and dynamic resource scheduling unit; The ISP and the NPU perform bidirectional data interaction through a direct data channel instead of shared memory; The intra-frame cooperative controller is used for dynamically delimiting an adaptive region of interest (ROI) in the image according to the feature position information fed back by the NPU, and simultaneously transmitting only the preprocessed feature data of the adaptive region of interest (ROI) to the NPU through the direct data channel; The inter-frame optimization module is used for identifying and transmitting only the visual feature reservation data of the dynamic change region according to the analysis of the difference image between consecutive frames, and actively adjusting the imaging parameters of the ISP in the next frame in combination with the inter-frame trend prediction signal output by the NPU; The dynamic resource scheduling unit is used for monitoring the load states of the ISP and the NPU in real time, and dynamically migrating the computing resources or storage bandwidth between the ISP and the NPU when a preset triggering condition is met.

2. The ISP and NPU based intelligent vision chip fusion architecture according to claim 1, wherein, The delimiting rule of the adaptive region of interest (ROI) comprises: taking the fault point coordinates fed back by the NPU as the center, and delimiting the ROI region according to an adaptive radius, which is dynamically adjusted according to the device size. 3.The ISP and NPU based intelligent vision chip fusion architecture according to claim 1, wherein, The intra-frame cooperative controller is built-in with an environment task adaptation rule mapping table, which comprises: when a night vision scene is detected, the ISP enables 3D noise reduction and gain enhancement, and the NPU enables YOLOv5s-DeepSort model and sets the IoU threshold; when a backlight scene is detected, the ISP enables local exposure equalization and dynamic range expansion, and the NPU enables anti-halo feature enhancement identification mode; when a rain and fog scene is detected, the ISP enables transmittance estimation defogging algorithm, and the NPU enables fog robust feature extraction network. 4.The ISP and NPU based intelligent vision chip fusion architecture according to claim 1, wherein, The difference image analysis is performed by calculating the difference image of consecutive frames, and the adaptive double-threshold method is used for binarization of the difference image to separate the dynamic region. 5.The ISP and NPU based intelligent vision chip fusion architecture according to claim 1, wherein, The inter-frame optimization module adopts a visual feature reservation compression strategy based on structural similarity, specifically: for the pixel region in the neighborhood of the fault point with a gradient change rate exceeding a preset threshold, the original precision is forcibly reserved, and the remaining regions are transmitted by lossy compression. 6.The ISP and NPU based intelligent vision chip fusion architecture according to claim 1, wherein, The inter-frame trend prediction signal is the target state change trend between consecutive frames output by the NPU, and the ISP adjusts the exposure time, gain or white balance parameters in advance according to the trend using a linear prediction model, so as to suppress the brightness or contrast fluctuation of the next frame image.

7. The ISP and NPU based intelligent vision chip fusion architecture according to claim 1, wherein, The preset triggering condition of the dynamic resource scheduling unit comprises: when the ISP computing unit occupancy rate continuously exceeds a first threshold and the NPU task queue is empty, triggering temporary migration of part of the NPU computing unit to the ISP; when the NPU task queue length exceeds a second threshold and the ISP storage bandwidth idle rate is higher than a third threshold, triggering reuse of the storage bandwidth of the ISP to speed up the NPU data reading. 8.The ISP and NPU based intelligent vision chip fusion architecture according to claim 1, wherein, Further comprise: An integrated safety data bus is used for connecting the ISP, NPU, intra-frame cooperative controller, inter-frame optimization module and dynamic resource scheduling unit to realize data transmission. 9.The ISP and NPU based intelligent vision chip fusion architecture according to claim 8, characterized in that, Further comprise: The codec module and the security control unit, the integrated data bus directly transmits the image processed by the ISP to the codec module, and the analysis result of the NPU is uniformly output after being encrypted by the security control unit.

10. A construction method of an ISP and NPU-based intelligent vision chip fusion architecture, characterized in that, Comprise: A direct data channel is arranged between the ISP and the NPU, and the ISP and the NPU perform bidirectional data interaction through the direct data channel instead of shared memory; According to the characteristic position information fed back by the NPU, the intra-frame cooperative controller dynamically delimits the adaptive region of interest (ROI) in the image, and only transmits the preprocessed feature data of the ROI to the NPU through the direct data channel; According to the analysis of the difference image between the continuous frames, the inter-frame optimization module identifies and only transmits the visual feature reservation data of the dynamic change region, and combines the inter-frame trend prediction signal output by the NPU to actively adjust the imaging parameters of the next frame of the ISP; Through the dynamic resource scheduling unit, the load states of the ISP and the NPU are monitored in real time, and when the preset trigger condition is met, the computing resources or storage bandwidth between the ISP and the NPU are dynamically migrated.

Citation Information

Patent Citations

  • Programmable visual chip-based visual image processing system

    CN102665049B

  • High-definition camera video processing method and system

    CN117689577A

  • Visual image monitoring method and electronic chip

    CN120219330A