Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

50results about How to "Reduce video memory usage" patented technology

Underwater target detection method based on differential attention

The invention discloses an underwater target detection method based on differential attention. The method comprises the following steps: acquiring an underwater image to be detected; based on the underwater image data set, constructing a lightweight underwater target detection network model based on a differential attention mechanism; constructing a joint loss function, and performing end-to-end training on the lightweight underwater target detection network model to obtain a trained lightweight underwater target detection network model; performing mathematical equivalence fusion on the lightweight underwater target detection network model to generate a lightweight reasoning model; and inputting the obtained underwater image to be detected into the lightweight reasoning model to realize target detection of the underwater image. According to the method, a differential attention feature interaction mechanism is introduced, common-mode noise is stripped through differential calculation, and the detection precision and real-time performance of an underwater fuzzy scene and a tiny target are remarkably improved in combination with a feature pyramid network integrating dynamic scale perception and semantic gating.
Owner:DALIAN MARITIME UNIVERSITY

Segmented mixed reasoning method based on uncertain driving large language model

This invention relates to a segmented hybrid inference method for large language models based on uncertainty-driven approaches. The method includes acquiring current text data and historical state features to estimate the uncertainty index of the current segment; minimizing a unified scheduling objective function based on the uncertainty index to obtain a target inference pattern; performing inference calculations based on the target inference pattern to generate information contribution values ​​corresponding to key-value pairs; calculating the corresponding dynamic merging control probabilities based on the uncertainty index and information contribution values, and performing weighted merging or pruning on the key-value pairs to be merged to obtain compressed key-value pairs; defining a deviation metric and limiting the deviation metric to not exceed a preset upper bound determined by the dynamic merging control probability set and the uncertainty index; triggering a rollback process when the deviation exceeds this limit; otherwise, feeding back the compressed key-value pair state to the next segment for iterative iteration until the inference of all segments is completed; thereby reducing memory usage and inference latency while ensuring accuracy.
Owner:XIAMEN UNIV

Document tampering detection method and system based on image data processing

ActiveCN121810690BSolve the problem of feature insensitivityAchieve macrodynamic amplificationImage enhancementImage analysisComputer graphics (images)Algorithm
The present application relates to the field of digital image processing and information security, and discloses a document tampering detection method and system based on image data processing, comprising the following steps: first, extracting the noise residual and microscopic penetration characteristics of the document image, and constructing a physical potential energy field and a virtual viscous resistance field; then, using a Darcy law variant model for dynamic evolution, generating a virtual flow velocity vector field to simulate the sliding behavior of fluid in heterogeneous media; subsequently, constructing a heterogeneous graph based on the flow field divergence singular point and streamline trajectory, using a graph neural network to aggregate the node dynamics characteristics for deep reasoning, and finally generating a tampering positioning mask. The present application innovatively introduces fluid mechanics field theory, converts hidden static texture differences into significant dynamic flow field anomalies, solves the problem that the prior art is difficult to capture microscopic tampering traces, and significantly improves the detection accuracy and generalization ability in complex document scenarios.
Owner:DOROAD ENERGY CO LTD

Light-weight working face three-dimensional reconstruction method and device fusing geometric prior constraints

PendingCN122089938AOvercome the problem of image feature matching failureSuppress scale driftImage enhancementImage analysisPoint cloudAlgorithm
The invention relates to the technical field of computer vision and coal mine intelligent mining, and discloses a lightweight working face three-dimensional reconstruction method and device fusing geometric prior constraints, and the method comprises the steps: firstly screening a dynamic effective region based on a projection relation and texture features, and removing static redundant scenes; performing sparse visual reasoning and depth regression only for the effective area to generate an initial depth map; carrying out physical structure correction on the depth map by utilizing geometric priori such as vertical upright posts and parallel top plates of the hydraulic support; and finally, constructing an optimization objective function containing deviation constraints of the vertical column perpendicularity and the top plate parallelism, and completing high-precision fusion of the incremental point cloud. According to the method, inference delay is reduced by reducing invalid calculation, structural distortion and registration drift caused by an underground severe environment are effectively inhibited by using a geometric law, and real-time and high-fidelity reconstruction of a working face scene is realized.
Owner:CCTEG COAL MINING RES INST +1

Edge terminal-oriented memory efficient model fine tuning method and system

PendingCN121835815ASolving memory overrunResolving training interruption issuesCharacter and pattern recognitionInference methodsPathPingCollaborative intelligence
The invention discloses an edge terminal-oriented memory efficient model fine tuning method and system, which realize decoupling of a backbone network and a fine tuning process by constructing a parallel side network module, and avoid video memory occupation caused by backbone gradient return. A double-adapter module is arranged, modeling is carried out on same-layer features and cross-layer context information, and the feature expression precision is improved. And constructing a feature fusion module, and performing adaptive weighting on different path outputs through learnable gating. A backbone grouping module is integrated, dynamic grouping is carried out on a backbone network based on interlayer similarity, and redundant modules and calculation overhead are reduced. And through the fine tuning execution module, only back propagation and parameter updating are carried out on the side network, and a trunk freezing state is kept, so that the memory and training cost is reduced. According to the method, the memory efficiency and the training speed of the edge fine adjustment process are improved, high-precision model self-adaption is achieved with extremely low calculation burden, and application in the fields of cloud edge collaborative intelligence and the like of unmanned aerial vehicles, robots, vehicle-mounted terminals and the like is effectively supported.
Owner:HOHAI UNIV +1

Quantization method and device for realizing elastic KV cache by computing power through intelligent computing cloud platform

The application provides a method and device for quantifying elastic KV cache through computing power of an intelligent computing cloud platform, and relates to the technical fields of intelligent computing centers, intelligent computing centers, computing power infrastructure and intelligent computing cloud technology.The method comprises the following steps: S1, dividing historical tokens into multiple cache blocks, quantifying KV data and writing the data into corresponding cache blocks, and selecting multiple candidate anchor points; S2, calculating block-level summary data; S3, in response to a new target token, scoring the cache blocks to generate cache block scores and screening out candidate cache blocks; S4, calculating uncertainty index data and determining a target cache block with to-be-restored precision according to the uncertainty index data; and S5, locating an upstream target anchor point and locally playing back the historical tokens based on the upstream target anchor point to generate target high-precision KV data of the target cache block.The application can greatly improve the quantification effect of KV cache of the intelligent computing cloud platform.
Owner:DATACANVAS LTD

A method for rapid landslide area segmentation based on high-resolution satellite remote sensing data and deep learning algorithms

ActiveCN121962622Bavoid confusionAvoid redundant calculationsAccurate segmentationSegmentation system
This invention provides a rapid landslide area segmentation method based on high-resolution satellite remote sensing data and deep learning algorithms, belonging to the field of ground scene technology. The invention downsamples the original satellite remote sensing image to obtain a first image, extracts reflectance data from the red and near-infrared channels, calculates the reflectance difference and sum of reflectance values, and normalizes them to obtain a vegetation index to construct a vegetation layer. The first image and the vegetation layer are superimposed and stitched together, and feature enhancement is performed to obtain a vegetation feature map. This map is input into a convolutional neural network to output a landslide probability map. A two-dimensional coordinate set is constructed and cropped to obtain high-resolution test patches. The spatial gradient magnitude of the local vegetation index is calculated and weighted by a preset enhancement coefficient to generate edge-enhanced patches. The edge-enhanced patches are input into a semantic segmentation network to output a local landslide map, which is then backfilled into the initial image to obtain a landslide area segmentation map. This invention constructs a cascaded segmentation system to achieve efficient and accurate segmentation of landslide areas.
Owner:SOUTH CHINA UNIV OF TECH

A Method and System for Intelligence Generation Based on Feature Fingerprint Storage and Spatiotemporal Geometric Correction

PendingCN122313306ASmooth deploymentReduce video memory usageData streamEngineering
This invention discloses an intelligence generation method and system based on feature fingerprint storage and spatiotemporal geometric correction, belonging to the field of intelligent remote sensing image processing technology. The method involves: accessing heterogeneous data streams generated from multi-source remote sensing images in a synchronized time sequence; inputting a shared backbone network and a modality adaptation layer to generate a current feature stream in a unified semantic space; asynchronously retrieving historical baseline features from an in-memory feature fingerprint database based on geographic coordinates; fusing satellite imaging parameters with the current feature stream and inputting it into a spatial transformation network, resampling historical baseline features using a regression transformation matrix to achieve spatiotemporal geometric correction in the feature domain; weighted fusing of the current feature streams from each modality to generate a fused feature stream; and parallel interpretation and semantic encapsulation of the fused feature stream to generate a natural language intelligence report. This invention reduces memory usage and data throughput, shortens the intelligence generation cycle, reduces the false alarm rate, and achieves synergistic optimization of timeliness, accuracy, and automation.
Owner:XIAN HUIGUANG RIXIN OPTOELECTRONICS TECHNOLOGY CO LTD

Large language model segmented hybrid reasoning method based on uncertain driving

The invention relates to a large language model segmentation hybrid reasoning method based on uncertainty driving, which comprises the following steps of: acquiring current text data and historical state characteristics to pre-estimate an uncertainty index of a current segment; according to the uncertainty index, minimizing the unified scheduling target function to obtain a target reasoning mode; performing reasoning calculation according to the target reasoning mode, and generating information contribution degrees corresponding to the key values; according to the uncertainty index and the information contribution degree, calculating a corresponding dynamic merging control probability, and performing weighted merging or pruning on the to-be-merged key value pairs to obtain compressed key value pairs; defining a deviation metric, and limiting the deviation metric not to exceed a preset upper bound determined by the dynamic merging control probability set and the uncertainty index; if so, triggering fallback processing; otherwise, the compressed key value pair state is fed back to the next segment for loop iteration until reasoning of all segments is completed; therefore, the video memory occupation and the reasoning delay are reduced on the premise of ensuring the precision.
Owner:XIAMEN UNIV

A cloud multi-pre-training language model management and inference method and electronic equipment

This invention discloses a cloud-based multi-pretrained language model management and inference method and electronic device, including receiving model management requests and inference requests issued by tenants through a dispatcher; wherein, the model management request is specifically a tenant initiating a model management request to change the content and structure of the vBert model instance tree; a shallow feature lookup table is constructed and maintained through a manager to update the vBert model instance tree; and the model management requests and inference requests are scheduled and processed in a pipeline manner through a scheduler.
Owner:ZHEJIANG UNIV

A speech recognition fine-tuning method based on subspace decomposition and recombination

ActiveCN122575345Bachieve reorganizationImprove reasoning
The application discloses a speech recognition fine-tuning method based on subspace decomposition and reorganization. The low-rank space of the LoRA module is equivalently decomposed into multiple subspaces, and learnable weights are added to different subspaces, so that the reorganization of all subspaces is realized. The reorganized subspaces can be transformed into an efficient calculation mode during training and an inference mode during inference, which significantly improves the inference and storage efficiency, and solves the problem of high inference cost during high concurrency request while ensuring efficient fine-tuning of parameters.
Owner:HEBEI UNIV OF TECH

A cross-modal data processing system for safe operation of hydrogen refueling stations

PendingCN122286340AReduce labeling costshigh quality conversionData processing systemHandling system
This invention discloses a cross-modal data processing system for the safe operation of hydrogen refueling stations, including a raw data acquisition and processing module, a full-variable safety scanning module, a physical relationship coupling diagnosis module, an adaptive operating condition clustering module, a question-answer pair construction module, a hydrogen refueling station time-series command data acquisition module, and a fault type identification module. By constructing a three-layer semantic enhancement logic and model adaptation strategy, this invention effectively solves the technical problems in the prior art, such as the lack of supervision signals in the raw data of hydrogen refueling stations, the difficulty in identifying hidden faults under complex operating conditions, and the difficulty in adapting heterogeneous feature space models.
Owner:CHONGQING UNIV

A robot welding action generation method, system and electronic device

PendingCN122500689Aeffective guidanceReduce the amount of invalid calculationsAdaptive optimizationEngineering
The application relates to the technical field of intelligent manufacturing and computer vision, and discloses a robot welding action generation method and system and electronic equipment. The application realizes adaptive elimination of visual redundancy in the time sequence dimension through global time sequence merging, guarantees complete reservation of key visual detail features in a complex welding scene through diversity perception pruning, realizes collaborative simplification of text and visual features through attention weight screening pruning, and completes dynamic adaptive optimization of bimodal features through semantic perception pruning. The application solves the defect that the prior art cannot adaptively process visual features, removes redundant features layer by layer, reduces the calculation load, greatly improves the real-time performance of welding work, accurately reserves the core visual and text semantic features required for welding action generation, fully guarantees the control accuracy of the welding action, and significantly improves the adaptive perception capability and work reliability of the robot welding system in a complex scene.
Owner:ROOTCLOUD TECH CO LTD

A method and system for mongolian-chinese multi-expert neural machine translation based on decision layer relative strategy optimization

PendingCN122735745AImprove task adaptabilityImprove adaptability
The application discloses a Mongolian-Chinese multi-expert neural machine translation method and system based on decision layer relative strategy optimization, and the method comprises the following steps: acquiring Mongolian-Chinese bilingual parallel corpus and preprocessing to obtain Mongolian-Chinese translation training data; training a basic translation model based on the Mongolian-Chinese translation training data to obtain a Mongolian-Chinese task adaptation model; performing sensitivity analysis on the Mongolian-Chinese task adaptation model to obtain hierarchical sensitivity results of each network layer; according to the hierarchical sensitivity results and the complexity of an input sentence, different network layers of the Mongolian-Chinese task adaptation model are allocated with different quantization precisions to obtain a light-weight translation model; a course type training sequence is constructed, and an optimized Mongolian-Chinese translation model is obtained based on a composite reward function; a Mongolian sentence to be translated is input into the optimized Mongolian-Chinese translation model to generate a corresponding Chinese initial translation; and a final Chinese translation is determined according to the Mongolian sentence to be translated and the Chinese initial translation.
Owner:INNER MONGOLIA UNIV OF TECH

A multi-branch attention table model without decoder and a construction method and application thereof

This invention discloses a decoder-free multi-branch attention table large-scale model and its reinforcement learning training method. Addressing the problems of redundant decoder structures, high computational overhead, and mismatch between training objectives and evaluation metrics in traditional supervised learning models, this invention proposes a multi-branch attention network based on an encoder-only architecture. This method removes the decoder portion from the traditional Transformer architecture, directly extracting feature interaction information using a multi-branch encoder. During the training phase, the Group Relative Policy Optimization (GRPO) algorithm is introduced for reinforcement learning training, abandoning the Critic network in traditional RL and directly calculating the relative advantage by sampling a set of outputs from the same input. This invention significantly reduces the number of model parameters and memory usage, while improving the accuracy and inference speed of table data processing by directly optimizing sequence-level rewards.
Owner:SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD

Efficient hyperspectral remote sensing image generation method based on spatial-spectral information self-selection

The invention provides a high-efficiency hyperspectral remote sensing image generation method based on spatial spectrum information self-selection. The method mainly comprises the following three core steps: 1, multi-level self-selection spatial feature coding; 2, multi-scale semantic fusion based on mixed pooling; and 3, spatial prior guidance and waveband relevance modeling. Starting from the characteristics of a hyperspectral remote sensing image generation task, a generation framework guided from spatial characteristics to spectral information driving is constructed, and a spatial-spectral information self-selection mechanism and a lightweight calculation strategy are introduced, so that the model can adaptively and dynamically pay attention to important spatial regions and spectral bands. The problems of insufficient physical consistency, redundant feature information, high video memory occupation and the like of a current hyperspectral remote sensing image generation method are effectively relieved, and the method has the advantages of high reconstruction precision, low calculation complexity, high training reasoning speed and the like; and high-quality hyperspectral remote sensing image data support can be provided for remote sensing downstream tasks such as image classification and change detection.
Owner:BEIHANG UNIV

A 3D breast ABUS image classification method based on a tokenized bi-branch selective state-space model, electronic devices, and computer-readable storage media.

This invention discloses a three-dimensional breast ABUS image classification method, electronic device, and computer-readable storage medium based on a tokenized bi-branch selective state-space model. The method preprocesses the three-dimensional ABUS volume data and inputs it into a hierarchical pyramidal classification network. The network constructs local convolutional branches and a global state-space branch in its basic modules: the local branch uses lightweight three-dimensional grouped convolution to extract texture and boundary morphology; the global branch aggregates features into a token map through voxel token generation, then performs multi-axis bi-directional selective state-space scanning and adaptive fusion using routing weights. After detoxing, the core features are injected through a gating mechanism. Finally, the classification result is output through three-dimensional global average pooling and a fully connected layer, enhancing the ability to model long-range dependencies across slices with near-linear complexity.
Owner:HANGZHOU DIANZI UNIV +1

Large language model inference acceleration method for real-time voice interaction and electronic device

The application discloses a large language model inference acceleration method for real-time voice interaction and electronic equipment, and relates to the technical field of artificial intelligence. The method comprises the following steps: obtaining optimal recognition text output by an automatic speech recognition engine as a draft sequence; inputting input data containing the draft sequence into a target large language model, and calculating posterior probability distribution of all token positions in the draft sequence through single forward propagation parallel calculation; determining a model output candidate sequence according to the posterior probability distribution of each position, comparing the model output candidate sequence with the draft sequence, and completing inference output based on the comparison result. The application does not need to additionally train, deploy and maintain an independent draft model, effectively reducing system engineering complexity and memory occupation. Meanwhile, the application scheme is executed based on the original architecture of the target large language model throughout the whole process, without the need to modify the model network structure, and without the need to carry out retraining or fine-tuning, so that the original generality and recognition accuracy of the pre-trained large language model are completely retained.
Owner:ANHUI IFLYTEK UNIVERSAL LANGUAGE TECH CO LTD

A three-dimensional defect segmentation method, system and storage medium

ActiveCN122090071Aeasy to identifyDefect Boundary SmoothingCharacter and pattern recognition3D modellingComputation complexityThree dimensional morphology
This invention relates to a three-dimensional defect segmentation method, system, and storage medium, belonging to the interdisciplinary fields of computer vision, deep learning, and nondestructive testing. The three-dimensional defect segmentation result is obtained by inputting the industrial CT volume data of the object to be inspected into a 2.5D segmentation network composed of an encoder, a spatial feature displacement module, and a decoder. This invention achieves efficient three-dimensional context awareness, significantly improving defect segmentation accuracy and three-dimensional morphological integrity. The segmentation result has topological connectivity and smooth boundaries, and its computational complexity and memory usage are far lower than those of a full 3D network. It can meet the real-time requirements of industrial online inspection and has broad application prospects in high-end manufacturing.
Owner:SHENYANG RES INST OF FOUNDRY

GPU storage and calculation integrated pre-filling decoding separation large model reasoning acceleration method and system

PendingCN121960766AActivate the potential of parallel computingHigh load operationDigital data information retrievalInterprogram communicationZero paddingAlgorithm
The invention provides a GPU storage and calculation integrated pre-filling decoding separation large model reasoning acceleration method and system, and the method comprises the steps: obtaining an original text, carrying out the operation of the original text through a GPU, and generating a corresponding key value cache and a query vector; based on a storage and calculation integrated memory, respectively reading the key vector, the numerical vector and the query vector from the key value cache for calculation, and generating a context vector; the method comprises the following steps: acquiring a plurality of request sequences, performing dynamic scheduling based on a preset time window or a request quantity threshold value, and sorting according to sequence lengths; carrying out zero-filling packaging on the input sequence by utilizing a GPU (Graphic Processing Unit) and executing pre-filling calculation to obtain an initial key value cache and a query vector; and respectively conveying the key value cache and the query vector to a storage area and a ready queue.
Owner:SHANGHAI JIAOTONG UNIV

Visual information injection method and device based on multi-modal large model

The invention relates to the technical field of artificial intelligence, and provides a visual information injection method and device based on a multi-modal large model, and the method comprises the steps: inputting a visual object into a visual encoder in a visual fusion model, and obtaining a visual feature sequence outputted by the visual encoder; inputting the visual feature sequence into a visual steering rotation matrix generator corresponding to at least one layer of the large model in the visual fusion model to obtain a visual steering rotation matrix generated by the visual steering rotation matrix generator based on the learnable query vector sequence and the visual feature sequence; and determining a second parameter vector of the corresponding layer based on the visual steering rotation matrix and the first parameter vector of the corresponding layer in the large model, and realizing visual information injection by using the second parameter vector and the text feature sequence. On the premise that the length of the input sequence is not increased and the structure of the large model is not damaged, deep fusion of vision and language is achieved, and the efficiency and performance of the large model in multi-modal tasks such as image question answering and video understanding are improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

3D medical image one-step generative segmentation method and system based on average flow model and medium

PendingCN122244443AMeet real-time surgical navigationReduce computational overheadBiological modelsInference methods
This invention discloses a one-step generative segmentation method, system, and medium for 3D medical images based on the average flow model, belonging to the field of medical image processing technology. The invention acquires the 3D medical image to be segmented and anatomical condition information, samples Gaussian noise as an initial latent variable, and inputs it into a MeanFlow network after temporal embedding. Anatomical conditions are input into a VeloMod module to generate scale and offset tensors, and the MeanFlow network features are modulated pixel-by-pixel. The average velocity field is calculated using the average flow identity, and target distribution features are generated through one-step mapping. A 3D segmentation mask aligned with the original image space is output by a 3D decoder. This invention achieves single-step function evaluation and inference, significantly improving segmentation speed and anatomical fidelity. It possesses advantages such as small-sample generalization, multimodal robustness, missing modality compatibility, and strong interpretability, meeting the needs of real-time clinical navigation, intraoperative planning, and high-throughput screening. It has significant application value in the field of intelligent 3D medical image segmentation.
Owner:LANZHOU UNIV

Industrial model rendering method and device, medium and program product

The invention relates to the technical field of computers, and discloses an industrial model rendering method and device, a medium and a program product. The method comprises the steps of obtaining a primitive type corresponding to an industrial model, splitting the industrial model according to the primitive type, and obtaining each initial parameterized grid piece and corresponding control point data; performing three-cascade elimination on each initial parameterized grid piece according to the control point data through an amplification shader, obtaining a target parameterized grid piece, performing drawing instruction distribution on the target parameterized grid piece, and obtaining each target thread unit corresponding to the target parameterized grid piece and a drawing instruction corresponding to each target thread unit; and through each target thread unit, performing discrete processing on the target parameterized grid piece according to the corresponding drawing instruction based on the grid shader to obtain triangular grid data. The occupation of a video memory in industrial software can be reduced, refined elimination and infinite precision display of a single entity are realized, and the utilization rate of GPU computing power by the industrial software can be improved.
Owner:BEIJING GLORY PKPM TECH CO LTD

An Efficient Fine-Tuning Method and System for Visual Transformer Parameters Based on Side-Rank Tuners

PendingCN122676477AReduce video memory usageAchieve multi-granularity collaborative adaptation
This application relates to the field of autonomous driving perception technology, specifically a method and system for efficient fine-tuning of visual Transformer parameters based on side-level low-rank tuners. By inserting lightweight low-rank tuners in parallel on the sides of the frozen pre-trained backbone network, detection accuracy comparable to full-parameter fine-tuning is achieved with only a minimal number of trainable parameters. This significantly reduces GPU memory usage and storage overhead during training, making scenario-based adaptation of large models feasible on consumer-grade hardware. The frozen main path fully preserves pre-trained visual knowledge and achieves lossless transfer, while the side branches specifically learn task-specific transformation residuals. A zero-initialization strategy ensures strict alignment between the training starting point and the pre-trained state, effectively bridging the gap in cross-task feature space while fully protecting general knowledge. Multi-granularity collaborative adaptation of spatial interaction, channel transformation, and overall distribution is achieved.
Owner:HONEYCOMB (WUHAN) MICROSYSTEM TECH CO LTD

Steel surface defect detection method and device, computer equipment and storage medium

The invention relates to a steel surface defect detection method and device, computer equipment and a storage medium, and belongs to the technical field of steel surface defect detection. The method is improved based on a YOLOv10 model, and comprises the following steps: embedding an ECA attention mechanism in a backbone network C2f module, and enhancing key feature representation; dEConv details are introduced into the convolution layer to enhance convolution, and defect space details are supplemented; the detection head is integrated with an MBConv module, and the positioning precision and the detection speed are balanced; the loss function is replaced by NWD, and the problem of small defect detection pain points is solved. According to the method, through multi-dimensional optimization of the model, the accuracy and robustness of steel surface defect detection are improved while light weight and high efficiency are kept.
Owner:ANHUI UNIVERSITY OF TECHNOLOGY

A hip joint ultrasound standard surface screening and measuring method and device based on key point detection and timing consistency

PendingCN122265153AEvenly screenedEasy to verify directlyMedical data miningImage analysisHeat mapRadiology
The application relates to a hip joint ultrasound standard surface screening and measuring method and device based on key point detection and time sequence consistency, and the method comprises the following steps: acquiring a hip joint ultrasound image sequence; detecting a plurality of anatomical key points in the ultrasound image, outputting a multi-channel response heat map, and obtaining anatomical key point coordinates; constructing a comprehensive consistency scoring function based on the anatomical key points and the time sequence change of adjacent frames of the hip joint ultrasound image, and calculating a consistency score; when the consistency score exceeds a preset threshold, the corresponding frame is determined as a candidate standard section frame; the development state of the candidate standard section frame is reviewed and distinguished to obtain a development state classification result; the alpha angle and the beta angle are calculated, the angle line and the key points are superimposed and displayed on an original drawing, a visual result is generated, and a structured report containing the visual result, the anatomical key point coordinates, the angle parameters and the development state classification result is output. Compared with the prior art, the application significantly improves the automation degree and the result interpretability while maintaining high-precision measurement.
Owner:SHANGHAI UNIV

A traffic flow prediction method and device based on large language model semantic enhancement

The application discloses a traffic flow prediction method and device based on large language model semantic enhancement, and belongs to the technical field of intelligent transportation systems. The method first acquires and pre-processes historical monitoring data of road network traffic sensors, and then inputs a pre-trained model to realize flow prediction. During model training, based on historical traffic space-time sequences and sensor distance matrices, statistical portraits such as sensor daily variation coefficients, information entropy and morning and evening peak ratios are extracted, semantic embedding vectors are generated by a large language model, and a global semantic similarity graph is constructed. After data normalization, sliding window segmentation is used to construct samples and divide training set, validation set and test set. The model is composed of dynamic adaptive fusion gate and lightweight space-time block, the training set is used to learn parameters, the validation set is used to supervise training and select the optimal weight, and finally the performance is verified in the test set. The method improves the traffic flow prediction accuracy and generalization by mining node correlation through semantic enhancement and combining lightweight space-time modeling.
Owner:HUAQIAO UNIVERSITY +1

Edge terminal intelligent model compression acceleration method

PendingCN122088576AHigh precisionSolve the waste of computing powerResource allocationBiological modelsData packElectrical battery
This invention provides a method for accelerating the compression of intelligent models on edge terminals, relating to the field of edge intelligence. Its key feature is that it includes: acquiring real-time operating status data of the edge terminal and feature information of the current input data, wherein the real-time operating status data includes computing resource load, memory occupancy, battery level, and device temperature. The advantages of this invention are: by sensing the device status and input complexity in real time, and utilizing reinforcement learning to dynamically generate an optimal strategy combination including pruning, hybrid quantization, and early termination mechanisms, this method overcomes the limitations of traditional static compression, achieving a dynamic balance between computing power, accuracy, and energy consumption. It not only significantly reduces inference latency and memory usage but also adaptively adjusts based on battery level and temperature, effectively solving the problems of difficult deployment, high heat generation, and short battery life of large models on resource-constrained edge devices, and greatly improving the real-time performance and stability of intelligent applications.
Owner:陈世恩

Wafer scanning image plane visualization dynamic loading rendering method and system

ActiveCN122547303BAvoid full loadingReduce memory
The application discloses a wafer scanning image plane visualization dynamic loading rendering method and system, and the system comprises a layout analysis and parameter initialization module, a multi-thread index construction module, a panoramic bottom layer static rendering module, a viewport adaptive switching module, a dynamic resource scheduling module and a cyclic rendering updating module. The method realizes automatic switching rendering of a thumbnail and a high-definition image by constructing a global layered index library and a panoramic thumbnail base map, and comparing the number of images in a screen viewport with a preset threshold in real time. Meanwhile, in combination with a preloading mechanism based on operation vector prediction and a two-stage hierarchical release strategy of images outside the viewport, high-definition instant display of local microscopic defects is realized while ensuring smooth panoramic browsing. The application greatly reduces system resource occupation, and realizes non-jitter visualization interaction of about 500,000 wafer images of 10 TB scale.
Owner:GUANGDONG SOLUDA TECHNOLOGY CO LTD

Cross-view geographic positioning method based on hybrid vision Mama network and centripetal block scanning

The invention discloses a cross-view geographic positioning method based on a hybrid visual Mama network and centripetal block scanning, and the method comprises the steps: constructing a hybrid backbone network comprising four stages, employing a convolutional neural network to extract robust local structure features in a shallow layer, and introducing a visual Mama architecture in a high layer to capture long-distance dependence and global context; a centripetal block scanning strategy (CBS) is designed according to the prior characteristic that an unmanned aerial vehicle image target is usually located in the center of a view, a feature map is divided into nine squares, sequence modeling is carried out in a sequence from outside to inside, and feature aggregation of the center target is enhanced while local continuity is kept; an InfoNCE loss function is used for training, and the feature space distance is optimized. According to the method, the calculation complexity is reduced, and meanwhile, the accuracy and robustness of cross-view geographic positioning are remarkably improved.
Owner:NANJING UNIV OF SCI & TECH