Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2162results about "Processor architectures/configuration" patented technology

Direct3D memory model compatible method based on adaptive occupied resources

The invention discloses a Direct3D memory model compatible method based on adaptive placeholder resources, which comprises the following steps: establishing resource metadata, accessing a scene library, checking an instruction template and double sandboxes when a DXVK is started, describing related resources by the DXVK through the metadata after a D3D application is started, completing resource mapping and metadata dynamic updating, and allocating the resources to the corresponding sandboxes; the DXVK compiles an application shader code, identifies a resource access instruction, matches an access scene and a check template by combining a pipeline stage and binding slot query metadata, instantiates the check instruction and adds the check instruction to the front of a target instruction; according to the unbound resources, adaptive generation of corresponding occupied resources is carried out, the authority and the state are configured, the unbound resources are bound to a Vulkan descriptor set to replace VKNULLHANDLE, and adaptation of the shader interface and the pipeline state parameters is completed to create PSO (Particle Swarm Optimization); and intercepting an application rendering instruction analysis parameter to construct a command buffer area, submitting the command buffer area to a Vulkan queue to trigger a GPU (Graphic Processing Unit) to execute rendering, and finally completing complete conversion and adaptation from the D3D rendering logic to the Vulkan.
Owner:北京麟卓信息科技有限公司

Direct3D rendering model compatible method based on dynamic template pool

The invention discloses a Direct3D rendering model compatible method based on a dynamic template pool, which comprises the following steps: establishing three types of mapping tables between D3D and Vulkan when compiling DXVK, constructing resource metadata, and creating a core, extended and temporary three-level template pool according to the mapping tables after starting; when the D3D application creates resources, metadata is initialized, parameters are verified, physical memories are allocated and grouped, resource groups are pre-verified, a batch binding command is generated, memory binding is completed, and the metadata is updated; when a resource view is created, a template is matched from a template pool, a handle is generated after instantiation, a resource handle is associated, and a descriptor is bound; when a rendering state is set and a rendering instruction is executed, the rendering state and the rendering instruction are respectively converted into a Vulkan related state and a Vulkan related instruction through a mapping table, a handle is bound after a PSO cache is inquired, a command buffer area is submitted to a GPU queue to execute drawing, and compatible operation of the D3D application on a platform supporting a Vulkan operating system is realized under the condition that the GPU does not support VKKHRmaintence5 and VKKHRmaintence6 extension.
Owner:北京麟卓信息科技有限公司

Graph rendering system and method based on OpenGL SC and related product

The invention discloses a graphic rendering system and method based on OpenGL SC and a related product. The system comprises a rendering interface module used for providing a high-order graphic rendering function interface for an application program to receive graphic modeling data; the data conversion module is used for converting the graphic modeling data into intermediate drawing data conforming to the OpenGL SC standard; the graphic processing module is used for executing graphic calculation and rendering logic based on the intermediate drawing data and generating a rendering instruction; and the graph drawing module is used for calling the underlying OpenGL SC function library to execute the rendering instruction so as to complete graph rendering. According to the graphic rendering system, the underlying graphic pipeline is packaged inside, the development threshold and complexity are reduced through the high-order graphic interface, and the development efficiency and reliability are improved. All the modules are standardized and coordinated to realize complete process packaging from data to rendering, so that optimization and maintenance are integrated inside, and the maintainability, expandability and stability of the system are enhanced.
Owner:CHINA TECHENERGY

Coordinating processing tasks between one-dimensional processing engines and two-dimensional processing engines

In various examples, systems and methods are disclosed relating to coordinating and synchronizing the actions of different types of processors with low latency. Different types of processors may perform better at different types of tasks. By coordinating the processing of a one-dimensional processor such as a vector processing unit (VPU) and the processing of a two-dimensional processor such as a pixel processing engine (PPE), an overall speed of task completion can be improved.
Owner:NVIDIA CORP

Memory shader

The invention relates to a memory shader. The programmable atomic memory shader execution circuitry is a seamless component of the hierarchical memory system, and can receive and execute programmable atomic operation calls from any number of processors. Programmable atomic memory shader execution circuitry close to memory allows execution circuitry to access shader programs stored in memory, thereby eliminating latency that may occur when an upstream processor exchanges shader instructions, data, and memory lock / unlock commands with the execution circuitry. The locked / unlocked state of the programmable atomic memory shader execution circuitry (e.g., in the L2 or L3 cache memory) allows the system to quickly lock the memory resources, perform one or more operations in an atomic manner over one or more cycles, and then quickly unlock the memory resources.
Owner:NVIDIA CORP

Proximity-based generation of surface textures for simulated environmental systems and applications

In various examples, a simulation platform generates a simulated driving environment by processing road map data to derive the location of wear-related visual artifacts for sections of a lane surface. Using map data, the simulation platform generates texture maps for aesthetic road rendering, which are used to apply textures to a 3D polygon topology mesh. The simulation platform generates visual artifacts representing the wear and tear of the road surface based on a calculation of one or more lane feature distances derived from the map data. The simulation platform calculates distances associated with reference line data derived from an image to render a texture of one or more lane features.The spacing is used to determine how the appearance of the texel is adjusted to include wear-related visual artifacts when displayed on a lane of the simulated driving environment.
Owner:NVIDIA CORP

Medical image segmentation method based on lightweight visual basic model

PendingCN121837632AImage enhancementImage analysisBoundary precisionFeature extraction
The invention discloses a medical image segmentation method based on a lightweight visual basis model, and the method comprises the steps: achieving the cross-modal and cross-anatomical region universal feature extraction through introducing a multi-scale token aggregation and a lightweight decoder; interlayer tokens are fused in a layered mode, global semantics and local textures are considered, and small target and boundary precision is remarkably improved. According to the multi-stage field adaptive pre-training, firstly, a fine-grained structure is captured by self-distillation, reconstruction and regularization composite loss, then, teacher-student distribution is aligned by Gram matrix anchoring loss, and finally, field migration and representation enhancement under the non-labeling condition are realized through high-resolution data refining. The lightweight compression selectively removes part of self-attention and retains MLP, and reduces parameter quantity and calculation quantity on the premise of almost no precision loss, so that the model can perform real-time reasoning on edge equipment. The end-to-end process reduces the annotation dependence, shortens the fine adjustment period, can support multi-organ and multi-focus segmentation through a unified frame, and improves the clinical deployment efficiency and generalization ability.
Owner:NANJING HEIKE ZHINING MEDICAL EQUIPMENT CO LTD

Methods and systems for tensor network contraction based on local optimization of contraction tree

Methods and systems for tensor network contraction are provided. A method implemented by a computing host comprises obtaining a contraction tree associated with a tensor network, wherein a plurality of vertices and edges of the contraction tree correspond to a set of tensor nodes and indices of the tensor network, respectively; iteratively performing operations until a termination condition is satisfied, the operations including selecting a sub-graph of the contraction tree; replacing the sub-graph with a local optimal sub-graph; and obtaining an optimized contraction tree including the local optimal sub-graph; and outputting the optimized contraction tree.
Owner:ALIBABA GROUP HOLDING LTD

Intelligent grading and excess subscription management system and method for GPU video memory

The invention provides an intelligent grading and excess subscription management system and method for a GPU video memory, and relates to the technical field of GPU video memory management, and the method comprises the steps: firstly building a grading storage system comprising a first performance region and a second performance region; then receiving a video memory allocation request containing task priority and service quality requirements, and monitoring a data access mode and access frequency; the initial placement position of the data block is dynamically determined through the intelligent engine according to the task priority, the data access mode and the frequency; predicting target access probability distribution of each data block by adopting an LSTM network; then migrating data between the first performance area and the second performance area through a swap-in and swap-out mechanism based on the distribution and periodicity characteristics, and ensuring the service quality of the first performance area; finally, the global oversale proportion is dynamically adjusted through an overpurchase safety management module, an independent virtual address space is distributed for each task, and safe overpurchase can be achieved through the process.
Owner:HANHOU (BEIJING) TECH CO LTD

Thread group scheduling method and device for GPU (Graphics Processing Unit), graphics processing unit and equipment

The invention discloses a thread group scheduling method and device for a GPU (Graphics Processing Unit), the GPU and equipment. The method comprises the following steps of: 1) receiving a scheduling request of a thread group; 2) task type priority scheduling; according to a thread group weight value set by a user, obtaining execution priorities of the vertex thread group and the fragment thread group in the current scheduling period; 3) instruction type priority scheduling; pre-analyzing to-be-executed instructions of the thread group, and determining a priority sequence of schedulable instruction types; and 4) priority scheduling of the thread groups: selecting the thread group with the highest priority for scheduling according to the specified task type priority and instruction type priority in combination with the thread group generation time. The invention provides a thread group three-level scheduling strategy so as to improve the instruction throughput rate and the key task response speed of the GPU under the complex load.
Owner:WUHAN LINGJIU MICROELECTRONICS CO LTD

Method and system for scheduling heterogeneous GPU (Graphics Processing Unit) by cross-cluster management software

The invention provides a method and system for dispatching heterogeneous GPUs through cross-cluster management software, and relates to the technical field of GPU resource dispatching.The method comprises the steps that physical topology data, hardware specification data and historical task records of GPU nodes are obtained through the cross-cluster management software, an association degree matrix between GPU units is established, and a topology grouping set is generated; calculating a communication efficiency index based on topology detection analysis, constructing a topological graph model, and determining an optimal communication path of a multi-card task; generating a perception scheduling strategy in combination with a multi-objective optimization algorithm, and selecting an optimal GPU combination from the grouping set; in the task execution process, a scheduling strategy is dynamically adjusted according to the real-time cluster state and task efficiency data, continuous adaptation of a physical topological structure and task requirements is kept, and finally efficient collaborative scheduling of cross-cluster heterogeneous GPU resources is achieved. According to the method, the scheduling efficiency of heterogeneous GPU resources in a distributed computing scene and the overall performance of the system are improved.
Owner:BEIJING HUAHENG SHENGSHI TECH CO LTD

Vehicle-mounted image generation device, in-vehicle infotainment system and vehicle

The invention relates to the technical field of vehicles, and discloses a vehicle-mounted image generation device, a vehicle machine system and a vehicle. The method comprises the following steps: at least acquiring first point cloud data and first image data through a data fusion module, and converting the two data into a bird's-eye view feature map through a preset conversion algorithm; performing semantic analysis on the user voice through a preset semantic analysis model by using a voice analysis module; generating a vehicle-mounted application image at least applied to head-up display and / or screen display according to the bird's-eye view feature map and the semantic analysis result by using an image generation engine through a preset compression model; and terminating the generation process of the vehicle-mounted application image in a preset time period after the automatic emergency braking signal is received through the emergency interruption module, and at least applying the graphic processing computing power of the generation process to a vehicle collision avoidance decision. Therefore, the problem that in an existing vehicle-mounted image generation technology, due to the fact that the computing power is too large, vehicle decision making is not timely under the emergency situation can be solved, and the driving safety of a user can be guaranteed.
Owner:CHINA FAW CO LTD +1

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Task processing device, task processing method, electronic equipment and storage medium

The embodiment of the invention provides a task processing device, a task processing method, electronic equipment and a storage medium. The task processing device comprises a calculation scheduling core particle and a function core particle which are integrated in a single package, the calculation scheduling core particle comprises a hardware subsystem and a calculation core, and the hardware subsystem is configured to receive and analyze task information to obtain at least one task. And assigning at least one task to at least one of the functional core and the computing core according to the task type; the computing core is configured to execute computing operations according to tasks allocated by the hardware subsystem, and the functional core particles comprise at least one of a quantum computing management core particle and a ray tracing core particle. According to the task processing device, through a unified task scheduling distribution mechanism, the adaptive capacity of the task processing device in an end-side mixed task scene is remarkably improved, so that a single computing chip can efficiently support various heterogeneous workloads.
Owner:SHANGHAI BIREN TECH CO LTD

RDMA queue pair multiplexing method based on shared memory

An RDMA queue pair multiplexing method based on a shared memory efficiently maps a large number of logical queue pairs (LQPs) to a small number of physical queue pairs (PQPs), and combines shared memory zero-copy communication and O (1) active queue management, so that ten thousand-level LQPs are supported by a small number of PQPs, the overhead of a single-connection memory is reduced by about 60%-64%, and context jitter of a network interface card (NIC) is avoided; the connection establishment time delay and the short flow end-to-end time delay are greatly reduced; through O (1) active tracking and batch processing, CPU idling and system calling overhead are greatly reduced, and high throughput and expandability are achieved; the method can be implemented in a pure user mode, a kernel or NIC does not need to be modified, and the method can be immediately deployed in an existing RDMA environment.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Pprimitive set distribution method and device, computer program product, equipment and non-transitory storage medium

The embodiment of the invention relates to a primitive set distribution method and device, a computer program product, equipment and a non-temporary storage medium. The method comprises the steps that a screen is divided into N blocks; mapping the primitive set and the N blocks one by one to obtain N mapping results; distributing effective mapping results in the N mapping results to M processing paths; wherein the mapping result refers to output information generated after at least one target primitive in the primitive set and at least one target block in the N blocks are mapped; wherein N and M are integers greater than or equal to 1. Therefore, data distribution can be flexibly performed according to the mapping results of the primitive set and the blocks, only the effective mapping results are distributed to each compression processing path, and the resource utilization rate of each path and the compression processing speed of the screen mapping data are improved.
Owner:MOORE THREADS TECH CO LTD

Urinary system tumor big data analysis system

The invention relates to the technical field of medical data analysis, in particular to a urinary system tumor big data analysis system which comprises a target kernel generation module, a kernel matrix calculation module, a parameter optimization module and a risk division module. According to the method, a target kernel matrix reflecting clinical prognosis differences is constructed and serves as an optimization reference, a radiomics and genomics feature kernel matrix is generated through hardware acceleration parallel computing, kernel function width parameters are dynamically iteratively updated based on an alignment degree numerical value so as to ensure that multi-modal feature distribution is highly matched with a prognosis label, and the accuracy of the multi-modal feature distribution is improved. A multi-dimensional feature space containing rich pathological information is constructed by combining a weighted fusion mechanism after centralization processing, so that a support vector machine is trained to determine a high-robustness decision boundary, and precise division of tumor risk levels is realized while high-dimensional data calculation delay is greatly reduced; and the reliability and timeliness of auxiliary diagnosis and treatment results in a complex pathological environment are effectively improved.
Owner:FIRST AFFILIATED HOSPITAL OF DALIAN MEDICAL UNIV

Loading storage circuit and graphics processor

The invention provides a loading storage circuit and a graphics processor, and relates to the technical field of graphics processing. The loading storage circuit comprises an instruction scheduling module, an address generation module and a data service module, and is provided with a buffer area module which comprises an operand buffer area and an effective data buffer area and is used for caching information required by address calculation and data access; the instruction scheduling module is used for collecting a data access instruction and outputting instruction information; the address generation module calculates a target access address according to the instruction information and the operand; and the data service module executes corresponding data loading or storage operation on the target cache unit. According to the scheme, the operands and the valid data are pre-cached, so that overflow and stagnation caused by inconsistent processing rhythms among modules can be avoided, and the parallel processing capability of an assembly line is improved; through the independent address calculation and data access process, the stability of the access time sequence and the data processing efficiency can be improved.
Owner:MOORE THREADS TECH CO LTD

Surface defect detection method for copper-aluminum composite board

The invention relates to the technical field of industrial visual inspection, and discloses a surface defect detection method for a copper-aluminum composite board. The method comprises the following steps: acquiring an original plate surface image through image acquisition equipment, and preprocessing to obtain a standardized image; performing texture field decomposition on the standardized image, separating out independent texture components of the copper layer and the aluminum layer, and calculating a coherence graph among the components; inputting the texture component and the coherence graph into a composite structure discrimination network, guiding feature fusion by the network according to the coherence graph, performing anomaly judgment, and outputting a pixel-level initial defect graph; and performing morphological and topological rule processing on the initial defect graph, generating a final defect contour, and performing superposition labeling on the final defect contour and the original graph. According to the method, the single material layer defect and the composite interface defect can be distinguished by separating the texture of the material layer and utilizing coherence guide analysis, and the detection precision is improved.
Owner:SHENZHEN TENGXIN PRECISION ADHESIVE PROD CO LTD

Proximity-based surface texture generation for simulated environment systems and applications

Proximity-based surface texture generation for simulated environment systems and applications is disclosed. In various examples, a simulation platform generates a simulated driving environment by processing road map data to infer locations of wear-related visual artifacts for portions of a roadway surface. The simulation platform uses the map data to generate texture maps for aesthetic road rendering, which are used to apply texture to the 3D polygon topology mesh. The simulation platform generates visual artifacts representing use and wear of the roadway surface based on calculating one or more distances from roadway lane features derived from the map data. The simulation platform calculates a distance from the one or more roadway lane features associated with reference line data derived from the image to render the texture. The distances are used to determine how to adjust the appearance of texels to include wear-related visual artifacts when rendering wear-related virtual artifacts on the roadway of the simulated driving environment.
Owner:NVIDIA CORP

Cooperative parallel memory allocation

Apparatuses, systems, and techniques to perform multi-threaded memory allocation in parallel by one or more software programs being performed on a parallel processing unit (PPU), such as a graphics processing unit (GPU), or any other processing unit capable of supporting multi-threaded software execution. In at least one embodiment, one or more software programs expressed in part by code using an application programming interface for parallel computing, such as CUDA, perform allocation, search, and deallocation of memory efficiently and in parallel on a GPU.
Owner:NVIDIA CORP

Direct3D 12 sampler addressing mode compatible method based on double buffer preprocessing

The invention discloses a Direct3D12 sampler addressing mode compatible method based on double buffer preprocessing, which comprises the following steps of: establishing a preprocessing parameter cache pool, a first association table, a second association table and a resource management linked list in a system initialization and starting stage, and registering a Vulkan expansion structure body for expanding a texture coordinate range and marking MSAA sampling point offset; when the textures are created, the VKD3D judges whether the textures are MIRROONCE related textures or not, allocates resource marks to the related textures, sets dependency marks, allocates double buffers and expands a coordinate range, and executes standard conversion if the textures are not related; when a sampler is created, marking an analog sampler, constructing and configuring a Vulkan sampler structural body, and updating an association table and a resource management linked list; when a pipeline is created, the D3D12 shader is converted into SPIR-V to update the second association table; and when the recording command list is applied, generating a command buffer, and submitting the command buffer to complete rendering.
Owner:北京麟卓信息科技有限公司

Deep learning system

A machine learning system is provided to enhance various aspects of machine learning models. In some aspects, a substantially photorealistic three-dimensional (3D) graphical model of an object is accessed and a set of training images of the 3D graphical mode are generated, the set of training images generated to add imperfections and degrade photorealistic quality of the training images. The set of training images are provided as training data to train an artificial neural network.
Owner:MOVIDIUS LTD

Parallel connected domain analysis method based on neighborhood labeling, computer equipment and storage medium

The invention relates to the technical field of image processing, and discloses a parallelization connected domain analysis method based on neighborhood labeling, computer equipment and a storage medium, through high parallelization design, the steps of father node, root node and area junction point searching, mapping establishing, re-numbering and the like are all designed into an independent parallel computing mode, and the parallel computing mode is designed into a parallel computing mode. According to the method, the acceleration capability of parallel computing hardware such as a GPU can be fully utilized, the processing speed is remarkably improved, the method is easy to expand, large-scale high-resolution images can be efficiently processed, the area of each connected domain is counted by adopting a Hash optimization method of grouping local reduction and global merging, atomic conflicts are effectively reduced, the statistical efficiency is improved, and the calculation efficiency is improved. The robustness of the algorithm in a parallel computing environment is further enhanced, so that the efficiency of the connected domain analysis algorithm is greatly improved, rapid processing of high-resolution images is realized, and powerful support is provided for real-time performance and large-scale application in the field of computer vision and image processing.
Owner:GUANGDONG AOPUTE TECH CO LTD

Smart storage system

A vehicle includes multiple storage zones. A set of imaging sensors defines fields of view and each storage zone is included in at least one defined field of view. A controller is in communication with the set of imaging sensors and includes a set of processing modules configured to cooperatively implement a smart storage tracking system including a person recognition processing module and an object recognition processing module. The person recognition processing module includes a process for identifying unique persons using image analysis and the object recognition module includes a process for identifying unique objects using image analysis. The smart storage tracking system defines an object to person and location data entry in a database using the person recognition processing module and the object recognition processing module. The database is in communication with the controller and storing each defined object to person and location data entry.
Owner:GM GLOBAL TECHNOLOGY OPERATIONS LLC

Pixel result output method and device, processor, equipment, medium and product

The invention discloses a primitive result output method and device, a processor, equipment, a medium and a product, the primitive result output method is applied to a graphics processor, the graphics processor comprises at least two cores, and the at least two cores are parallel to subtasks of an original rendering task. The method comprises the steps that subtasks corresponding to the cores are executed in parallel through the two cores, and primitive results corresponding to the subtasks are output to temporary storage spaces corresponding to the cores; and copying the primitive result corresponding to each subtask to a primitive storage space in the temporary storage space corresponding to each core to obtain a target primitive result of the original rendering task.
Owner:MOORE THREADS TECH CO LTD

Sensing, storing and computing integrated bionic vision sensing module storing and computing method and system

The invention relates to the technical field of visual sensing modules, and discloses a sensing-storage-calculation integrated bionic visual sensing module in-storage calculation method and system.The method comprises the steps that photo-induced electrons excited by incident photons are directly injected into memristor oxygen vacancy conductive filaments through a shared electrode interface, and photocurrent driving signals are obtained; in a forward SET mode, driving memristor oxygen vacancy conductive filaments to form cross array conductivity distribution, and calculating row line driving voltage of a memristor cross array; applying the current to each cross point of cross array conductance distribution in a reverse RESET mode to obtain a flowing current of each cross point; and converging and summing the flowing current of each cross point at a column line node to obtain a column line output current, performing capacitor charging integration on the column line output current, and converting the column line output current into an analog domain convolution operation voltage, thereby avoiding intermediate multiple analog-to-digital and digital-to-analog conversion loss. And end-to-end low-delay and low-power-consumption processing from photon sensing to feature calculation is realized.
Owner:DONGGUAN TSIMSAFE ELECTRONICS TECH