Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

292 results about "Model compression" patented technology

Three-dimensional model compression transmission method and system based on dynamic feature perception

The invention relates to a three-dimensional model compression transmission method and system based on dynamic feature perception, and relates to the technical field of three-dimensional live-action modeling, and the method comprises the steps: obtaining the original three-dimensional data of a three-dimensional model, carrying out the multi-modal feature extraction based on a model scene, dividing a key region, a transition region and a non-key region, and generating a three-dimensional model partition map, compressing each region according to the partition map label and a preset compression algorithm to generate a multi-resolution LOD sequence, determining a transmission data hierarchy in combination with a network state and equipment parameters, and finally rendering the model according to the transmission data hierarchy, user behavior information and an environment state to obtain a target three-dimensional model. The technical effects of effectively compressing, transmitting and rendering the three-dimensional model according to various factors such as different area characteristics, network states, equipment parameters and user behaviors of the model and improving the transmission efficiency and the rendering quality are achieved.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Traffic accident intelligent detection system and method based on YOLOv12 improved architecture

The invention relates to a traffic accident intelligent detection system and method based on a YOLOv12 improved architecture. The system comprises a YOLOv12 enhanced feature extraction network, a multi-scale detection head, a time sequence information fusion module, a real-time reasoning optimization engine and an intelligent decision fusion system, and realizes collaborative optimization of local feature enhancement and global context modeling by constructing six core technology modules and adopting collaborative learning of a C2f-Attention mechanism and deformable convolution. According to the method, a composite loss function special for traffic accidents is innovatively designed, and adaptive fusion of multi-scale features and difficult sample mining are realized through a multi-objective optimization mechanism of Enhanced Focus Loss, IoU-aware Loss and Severage-aware Loss. According to the method, the problems of low detection precision and false alarm and missing alarm caused by illumination variation, shielding and motion blur in a traffic monitoring scene are effectively solved, in the test of an AccidentsDesection YOLOv8 data set, the mAP at 0.5 reaches 91.27% and is improved by 8.6% compared with that of YOLOv8, the reasoning speed reaches 67 FPS, experimental results show that the system has excellent performance in the aspects of detection precision, real-time performance and model compression, and the method is suitable for popularization and application. The method achieves a remarkable effect in traffic accident intelligent identification, and has a remarkable technical effect and industrial application value.
Owner:JIANGSU OCEAN UNIV +1

Multi-modal model compression and distillation method and system based on causal reasoning

The invention relates to the field of multi-modal neural network model compression, and particularly discloses a multi-modal model compression and distillation method and system based on causal reasoning, and the method comprises the steps: constructing a comprehensive causal discovery module, identifying a causal dependency relationship among the multi-modal features through information theory measurement, Granger causal analysis and intervention-based verification; executing an adaptive compression engine, and performing pruning, mixing precision quantification and low-rank decomposition based on a causal relationship; a cross-modal distiller is applied, and multiple loss function combinations are adopted to maintain the relationship between modals; and implementing a dynamic optimizer to carry out hardware perception and context-sensitive reasoning optimization. According to the method, the compression decision is guided through causal reasoning, the high compression rate is achieved while the key causal path is kept, and the deployment problem of the multi-modal model in the resource-constrained environment is effectively solved.
Owner:SHENZHEN UNIV

Method for detecting health state of household energy storage lithium battery

The invention relates to the technical field of energy storage lithium batteries, in particular to a method for detecting the health state of a household energy storage lithium battery. The method for detecting the state of health of the household energy storage lithium battery is characterized by comprising the following specific steps: S1, preprocessing charge-discharge cycle data; s2, performing multi-domain feature extraction; s3, an SA-XGBoost model oriented to static data and a TSO-LSTM-AM model oriented to dynamic data are constructed respectively, and SOH estimation is carried out on the SA-XGBoost model and the TSO-LSTM-AM model; and S4, finally, performing compression optimization on the model through knowledge distillation and quantification technologies to meet a lightweight deployment requirement. Compared with the prior art, in combination with historical charging and discharging parameters such as battery voltage, current and temperature, feature vectors are constructed through adaptive data cleaning and multi-domain feature extraction, the SOH estimation capability of static and dynamic data is optimized by adopting machine learning and a deep learning model respectively, and the deployment requirement of practical application is met by utilizing a model compression technology.
Owner:SHANGHAI PYTES ENERGY CO LTD

Multi-modal perception enhancement method for automobile cabin end side model

The invention provides an automobile cabin end side model multi-modal perception enhancement method, which comprises the following steps: S1) deploying an RGB camera in a cabin, and collecting facial biological characteristic data, hand interaction behavior data and cabin overall environment data of a driver; s2) adjusting image illumination and contrast by adopting an algorithm, dynamically focusing a key area through a target detection frame to process a shielding problem, and eliminating static redundant information; s3) constructing a neural network, extracting facial biological features and the like in parallel, and outputting low-dimensional feature vectors; s4) time sequence association is established, key area feature weights are enhanced through a space attention mechanism, and a logic relationship among different features is explicitly modeled; s5) aggregating the weighted feature maps by using feature attention pooling, and retaining a core feature channel in combination with a channel pruning technology to realize model compression; s6) dynamically adjusting the calculation priority and weight distribution of each feature extraction branch; and S7) the end side uploads the sample to the cloud side, and the cloud side generates an update package and pushes the update package to the end side to complete model iteration.
Owner:SAIC VOLKSWAGEN AUTOMOTIVE CO LTD

Miniaturized blood glucose intelligent monitoring and analysis method integrating edge calculation

The invention relates to the crossing field of computer technology and biomedical engineering, and discloses a miniaturized intelligent blood glucose monitoring and analyzing method integrating edge computing. The method comprises the steps that glucose concentration and time sequence multi-source physiological signals are collected through a wearable sensor; performing data segmentation and multi-modal feature alignment based on the circadian rhythm; inputting the features into a lightweight student model obtained through knowledge distillation for blood glucose prediction; and the reasoning frequency is adaptively adjusted in combination with a memory feedback mechanism. The system comprises a sensing unit, a preprocessing unit, a lightweight model execution unit and a risk early warning unit. While the prediction precision is guaranteed, the model is compressed to be within 150 thousand parameters, a 512 KB on-chip memory is adapted, and low-power-consumption and high-stability end-side continuous blood glucose monitoring and early warning are achieved.
Owner:FUZHOU STRAIT VOCATIONAL & TECH COLLEGE

Model compression and data enhancement fused lightweight deep counterfeit voice detection method

The invention provides a model compression and data enhancement fused lightweight deep forged voice detection method. The method comprises the steps of obtaining and processing a public voice data set and a large-scale self-supervision pre-training voice model; performing structured pruning and knowledge distillation to obtain a lightweight voice model; performing audio preprocessing and diversified data enhancement on the true and false voice samples to obtain an enhanced true and false voice data set; performing faking task joint fine tuning on the lightweight voice model to obtain a lightweight deep faking voice detection model; and locally deploying the model to obtain a localized counterfeit voice detection system, and carrying out real-time authenticity judgment. According to the method, calculation complexity and reasoning time delay are remarkably reduced, local deployment is carried out on a resource-limited end side platform, detection generalization and robustness are improved, and low-delay, low-power-consumption and high-robustness detection performance is achieved.
Owner:ZHEJIANG UNIV

Satellite-ground collaborative learning method and system and electronic equipment

The invention relates to a satellite-ground collaborative learning method and system and electronic equipment, and relates to the technical field of satellite remote sensing, and the method comprises the steps: compressing a basic model into a small model through a ground station, and enabling the small model to serve as a satellite in-orbit model; detecting the comprehensive distribution offset of the satellite during the in-orbit reasoning period, and when the comprehensive distribution offset is greater than a preset threshold value, sending related data to a ground station; calculating a weighted difference value between a prediction result of the on-orbit model and a prediction result of the basic model based on the related data, and triggering updating of the on-orbit model when the weighted difference value exceeds a preset threshold value; loRA fine tuning is carried out on the small model through the ground station so as to realize on-orbit model updating; and the steps are repeatedly executed until the updating frequency of the on-orbit model reaches the preset frequency, and the on-orbit model is aggregated into the basic model through the ground station to achieve model updating of the basic model. According to the scheme, the overall performance of the satellite on-orbit model facing the distributed external scene can be improved.
Owner:HANGZHOU INST FOR ADVANCED STUDY UCAS

Knowledge distillation-based 1D-CNN online partial discharge identification method

The invention belongs to the technical field of power equipment fault diagnosis, and particularly discloses a knowledge distillation-based 1D-CNN online partial discharge identification method, which comprises the following steps of performing Fourier transform on an original signal and then implementing four data enhancements of amplitude modulation, phase modulation, frequency band enhancement and frequency domain noise injection; constructing a multi-branch one-dimensional convolutional neural network to extract time-frequency features, and training a teacher model in combination with residual connection and an SE attention mechanism; compressing the teacher model into a lightweight student model by adopting a knowledge distillation strategy of a preset temperature parameter and a weight coefficient; the student model is converted into a mobile terminal compatible format, and a mobile phone APP is established through a QT upper computer; and online real-time identification is realized. According to the method, the problem of sample imbalance is effectively relieved, in addition, the knowledge distillation realizes model volume compression and reasoning speed increase while keeping relatively high accuracy, the calculated amount is small, the real-time performance is high, and the precision is high.
Owner:SHANDONG UNIV OF SCI & TECH

Deep neural network model optimization method based on hierarchical reinforcement learning and multi-agent collaborative distillation

The invention relates to the technical field of model lightweight, and particularly discloses a deep neural network model optimization method based on hierarchical reinforcement learning and multi-agent collaborative distillation, and the method comprises the steps: building a structured pruning searcher based on an ABC algorithm, constructing a pruning combination reduction strategy dynamic artificial bee colony pruning algorithm, and carrying out the optimization of a deep neural network model. Performing fitness evaluation to guide a search process, and outputting an optimal pruning network structure under resource constraint; establishing a staged distillation architecture and a multi-dimensional hierarchical loss function, and realizing smooth and progressive knowledge transmission between the teacher model and the assistant model; a fine-grained quantization scheme based on parameter classification is designed, differential bit widths are configured for weights, batch normalization parameters and activation output respectively, a quantization perception training loss function fusing a hardware delay look-up table model is constructed, hardware perception joint fine tuning of the network weights and quantization parameters is achieved, and the quantization precision of the network weights and the quantization parameters is improved. Therefore, the effect of remarkably improving the model compression efficiency on the premise of keeping the precision is achieved.
Owner:CHONGQING INST OF NEW ENE STOR MATER & EQUIP

A multi-modal model compression and distillation method and system based on causal reasoning

The application relates to the field of multi-modal neural network model compression, and specifically discloses a multi-modal model compression and distillation method and system based on causal reasoning, which comprises the following steps: constructing a comprehensive causal discovery module to identify the causal dependence relationship among multi-modal features through information theory measurement, Granger causality analysis and intervention-based verification; performing an adaptive compression engine to perform pruning, mixed precision quantization and low-rank decomposition based on the causal relationship; applying a cross-modal distiller to maintain the relationship among modes by using a variety of loss function combinations; and implementing a dynamic optimizer to perform hardware perception and context-sensitive reasoning optimization. The application guides the compression decision through causal reasoning, realizes high compression rate while maintaining key causal paths, and effectively solves the deployment problem of multi-modal models in a resource-limited environment.
Owner:SHENZHEN UNIV

Light-weight unmanned aerial vehicle infrared remote sensing target detection method and device based on knowledge distillation

The invention discloses a lightweight unmanned aerial vehicle infrared remote sensing target detection method and device based on knowledge distillation, and belongs to the technical field of infrared weak and small target image recognition and deep learning model compression. In order to solve the problems of low detection precision, high omission ratio, high calculation complexity and difficulty in deployment at edge equipment of an existing unmanned aerial vehicle infrared remote sensing image weak and small target detection algorithm, an infrared remote sensing image data set is constructed, and a basic network model containing an intra-scale feature efficient interaction and fusion module is constructed; a teacher model and a student model are designed, a lightweight student network is trained by using a channel dimension knowledge distillation mechanism, and optimization is performed by using a WIoU loss function based on a distance attention mechanism and a soft-NMS, so that multi-scale feature fusion and model compression are realized. According to the method, the accuracy and robustness of weak and small target detection are improved, the model parameter quantity and the calculation quantity are reduced, and the method is suitable for real-time application in the fields of environmental monitoring, military reconnaissance and the like.
Owner:GUANGDONG UNIV OF TECH

Neural network multi-objective optimization and FPGA hardware acceleration collaborative design method

The invention provides a neural network multi-objective optimization and FPGA hardware acceleration collaborative design method, and belongs to the field of deep learning model compression and hardware collaborative design. The method comprises the following steps: constructing a joint optimization space containing a neural network compression parameter and an FPGA hardware design parameter; a multi-target Bayesian optimization search strategy is adopted, iterative search is carried out in the joint optimization space, model precision, FPGA resource occupation and reasoning delay are synchronously optimized, and optimal candidate configuration is obtained; matching the compressed network structure with the FPGA parallel architecture by using a hardware-perceived pruning and quantification strategy; a multi-task performance prediction model is adopted to quickly predict the precision, resource occupation and delay of the optimal candidate configuration so as to accelerate the search process; according to the optimal configuration, a hardware accelerator code facing the target FPGA is automatically generated, and integration and implementation are completed. According to the method, collaborative optimization of neural network compression and hardware design is achieved, FPGA resource occupation can be remarkably reduced, the reasoning speed can be increased, and meanwhile the model precision is kept.
Owner:BEIJING JIAOTONG UNIV

Visual language large model compression and acceleration method based on redundant layer pruning

The invention belongs to the technical field of language image multi-modal fusion, and discloses a visual language large model compression and acceleration method based on redundant layer pruning, and the method comprises the steps: firstly obtaining the hidden state of visual and text lexical elements; calculating the importance score of each lexical element by integrating the internal and external attention of the modal; screening important lexical elements to locate a redundancy layer, and pruning according to a preset threshold value to obtain a pruned layer set and a reserved layer set; searching the nearest layer in the reservation set for each pruning layer, and calculating a feature difference matrix by using calibration data; performing singular value decomposition on the matrix to extract a low-rank subspace; and finally, using the weight of the subspace projection nearest retention layer to compensate the feature gap. According to the method, the model scale can be effectively compressed without training, the reasoning speed is improved, and the model performance is kept.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Multi-model compression deployment method and system based on pruning, medium and equipment

The invention discloses a multi-model compression deployment method and system based on pruning, a medium and equipment, and belongs to the technical field of artificial intelligence model compression and edge computation.The method comprises the steps that the network layer topological relation of each model to be deployed is extracted, and an internal dependency graph is constructed; iteratively calculating node importance scores, cutting off nodes and edges which have the lowest importance and are not on a critical path in each round, synchronously deleting a network layer and reconstructing a computational graph until the model volume reaches a target threshold value, and obtaining each first model; constructing a cross-model dependency graph of which the edge weight is determined by the data sharing rate and the interaction frequency by taking each first model as a node and the data flow direction between the models as an edge; and screening candidate edges with weights higher than a threshold value, and if the function overlapping degree of the corresponding model pair exceeds a preset threshold value, merging overlapped network components to complete model compression and deployment. By implementing the method, the problem that a plurality of models cannot be efficiently deployed at the edge end of the transformer substation due to limitation of computing power and memory in the prior art can be solved.
Owner:GUANGZHOU XINDILI ENERGY TECHNOLOGY CO LTD

Face recognition model training method, face recognition method and face recognition system

The invention provides a training method of a face recognition model, and a face recognition method and system, and relates to the technical field of face recognition, and the method comprises the steps: pre-training a teacher network; constructing a student network and establishing a knowledge distillation framework; dividing the feature vector output by the teacher network into blocks, and applying different weights according to the importance of each block so as to form weighted block mean square error loss; the cosine similarity of feature vectors output by the teacher network and the student network is used as a difficult sample indicator; the angle interval in additive angle interval loss is dynamically adjusted according to the indicator, and differential constraints are applied to samples with different difficulty levels. According to the method, the learning ability for key features is enhanced through weighted block mean square error loss, attention distribution for difficult and simple samples is optimized through dynamic additive angle interval loss, the identification precision, convergence speed and model compression effect of a student network are effectively improved, and the method is suitable for resource-limited edge device deployment.
Owner:JIANGNAN UNIV

Target detection method based on binary self-distillation Transform

The invention relates to a target detection method based on binary self-distillation Transform, and belongs to the technical field of model compression and target detection. The method comprises the following steps: inputting a training image into a constructed binary Transform student network for forward propagation, synchronously extracting full-precision QKV and binary QKV in the network, constructing a teacher feature similarity matrix and a student feature similarity matrix, and aligning feature distribution through value exchange to obtain feature self-distillation loss; performing dynamic smoothing processing on the target class labels, and performing weak supervision in combination with non-target class soft labels to obtain label self-distillation loss; and carrying out dynamic weighted fusion on the target detection loss, the feature self-distillation loss and the label self-distillation loss of the training image to obtain total loss, optimizing binary Transform student network parameters based on the total loss until the model is converged, and obtaining a binary target detection Transform model for target detection. The objective of the invention is to solve the technical problem that a full-precision model cannot be pre-trained in the prior art, thereby improving the target detection precision.
Owner:KUNMING UNIV OF SCI & TECH

AI chip reasoning acceleration method and system and medium

The invention relates to the related technical field of reasoning calculation, in particular to an AI chip reasoning acceleration method and system and a medium, and the method comprises the steps of collecting adaptive scheduling parameters, setting a hierarchical pipeline response gradient under a reasoning acceleration framework, loading the optimized scheduling parameters to a configuration register group, and setting reasoning self-driven circulation. The technical problems that load feature perception is one-sided, scheduling parameters are statically solidified, dynamic adjustment along with reasoning tasks cannot be achieved, adaptability to feature graph structure changes and activation mode updating in the reasoning process is poor, and reasoning stability is insufficient are solved, and a tracing strategy based on a three-dimensional load feature vector and attention guidance is achieved. The method has the technical effects that the matching degree of scheduling parameters and task states is improved, the scheduling parameters and model compression granularity are dynamically adjusted in combination with real-time energy consumption constraints, a federated on-chip learning architecture enables the parameter iteration updating efficiency of a layer sensitivity analysis unit to be improved, the reasoning stability and the continuous operation reliability are enhanced, and the method adapts to end side complex dynamic scene requirements.
Owner:TIANJIN RUILEX ENVIRONMENTAL PROTECTION ENG CO LTD

Compression of audio waveforms using neural networks and vector quantizers

This provides a method for compressing audio waveforms using machine learning models. [Solution] The method includes the steps of: receiving an audio waveform containing an audio sample for each of a plurality of time steps; processing the audio waveform using an encoder neural network to generate a plurality of feature vectors representing the audio waveform; and generating a coded representation of each of the plurality of feature vectors using a plurality of vector quantizers, each associated with a codebook of the code vectors. Each coded representation of each feature vector identifies a plurality of code vectors containing the respective code vectors from the codebooks of each vector quantizer. The method also includes the step of compressing each coded representation of the plurality of feature vectors to generate a compressed representation of the audio waveform.
Owner:GOOGLE LLC

Radiation particle identification method based on improved ResNet-18 network and CIS transient response technology

The invention discloses a radiation particle identification method based on an improved ResNet-18 network and a CIS transient response technology, and the method comprises the steps: collecting a CIS dark field image in a radiation environment, and generating a multi-modal data set according to neutron, proton and heavy ion radiation experiment sample images; a transient response image data set is obtained through theoretical calculation based on radiation analog simulation software; preprocessing the transient response image data set, and marking particle type, energy and angle information of each sample image as a label of multi-task learning; a pre-trained ResNet-18 model is used, an input layer is modified to adapt to the image size, and a channel attention module is added; and carrying out compression quantification on the model by adopting an optimizer in combination with gradient cutting through a self-adaptive category weight adjustment strategy. By constructing the intelligent feature extraction network, the system realizes automatic identification of transient response geometric features, and breaks through the bottleneck that a traditional method depends on artificially defined features.
Owner:YANGZHOU UNIV

Image recognition method, electronic device, readable storage medium and program product

The application discloses an image recognition method, an electronic device, a readable storage medium and a program product, relates to the technical field of artificial intelligence, and comprises the following steps: quantizing and pruning other layers except a target feature processing layer in a target branch structure of an image recognition network, and generating an equivalent virtual feature processing layer by performing structure transformation on multiple layers of the target branch structure; determining compensation parameters for compensating for quantization loss and pruning loss according to parameters of the virtual feature processing layer, and adjusting subsequent network layers to compress the target feature processing layer; and processing input images by using the compressed image recognition network to obtain image recognition results. The application can solve the problem of low recognition accuracy of related art multi-branch image recognition model compression, effectively compress the image recognition model of the multi-branch network structure, and reduce the required computing resources of the image recognition task while ensuring the image recognition accuracy.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Industrial detection methods, apparatuses, devices, products, and storage media

PendingCN122335705AVision inspectionAlgorithm
This application discloses an industrial inspection method, apparatus, equipment, product, and storage medium, relating to the field of industrial vision inspection technology. The method includes: responding to an industrial inspection command, acquiring an image to be inspected for membrane defects; processing the image using a preset defect detection model to obtain a defect detection result. The preset defect detection model is obtained by quantizing a target defect detection model, which is obtained by structural pruning and lightweight reconstruction of an initial defect detection model. The structural pruning and lightweight reconstruction are used to reduce the number of model parameters and computational load. This application reduces the number of model parameters and computational load through structured pruning and lightweight reconstruction, and then quantizes the model, enabling the quantized model to be deployed on edge devices while improving the model's inference efficiency, achieving efficient model compression and real-time inference.
Owner:SHENZHEN INSTITUTE OF INFORMATION TECHNOLOGY

Image information extraction method and device, electronic equipment and storage medium

The invention provides an image information extraction method and device, electronic equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: carrying out the model compression processing of an optical character recognition model, so as to obtain a lightweight optical character recognition model; performing text recognition on the input image by using the lightweight model at the edge equipment end to obtain text content and position information, then judging whether to trigger cloud processing according to a text recognition result, and sending the input image and the text recognition result to the cloud during triggering; the cloud server performs multi-modal information extraction through a large language model to obtain structured information, so that the problems that in the prior art, an optical character recognition model is difficult to adapt to the computing power of edge equipment, and end-cloud cooperation lacks a reasonable triggering mechanism, so that response delay or resource waste is caused, and the efficiency is high can be solved. And edge equipment cannot directly deploy a large model to realize high-precision structured information extraction.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Machine learning model compression

Techniques are described herein for a method of machine learning model compression. The method includes receiving a machine learning model comprising a plurality of blocks. The method further includes removing one or more blocks of the plurality of blocks to obtain an intermediate machine learning model comprising a subset of the plurality of blocks. The method further includes adding a block to the intermediate machine learning model to obtain a compressed machine learning model. The block generates an output corresponding to an output of the removed one or more blocks of the plurality of blocks. The method further includes executing the compressed machine learning model on a low resource device.
Owner:SALESFORCE INC

Large model compression method based on continuous layer pruning and endpoint tuning

The invention relates to a large model compression method based on continuous layer pruning and endpoint tuning, and the method comprises the steps: firstly introducing a learnable continuous interval soft mask, and building a differentiable hierarchical mask mechanism in a model in cooperation with a residual bypass; secondly, by minimizing the KL divergence between output distributions before and after pruning, the optimal pruning starting point and length are automatically learned, and adaptive selection of continuous layer segments is achieved; then, executing physical layer deletion according to the optimized interval parameters, and reconnecting the network structures before and after pruning; and finally, implementing an endpoint tuning strategy, only carrying out all-parameter fine tuning on key layers on two sides of the sheared interval, and recovering the model performance at the lowest calculation overhead. According to the method, through combination of differential interval search and end point directional optimization, accurate compression and high-performance maintenance of the depth dimension of the large model are realized, model storage occupation and reasoning delay are remarkably reduced, the model output reliability in a key task scene is guaranteed, and the method is suitable for large-scale popularization and application. The method is particularly suitable for efficient deployment of the large language model in a resource-constrained environment.
Owner:ZHEJIANG UNIV OF TECH

Neural network model compression method and device, equipment and medium

The embodiment of the invention provides a compression method and device of a neural network model, equipment and a medium, and the method specifically comprises the steps: generating a hybrid super network for a to-be-compressed neural network model; the hybrid hypernetwork includes: a searchable unit; the configuration parameters corresponding to the searchable unit comprise quantitative configuration parameters and network structure configuration parameters; sampling a quantitative configuration parameter and a network structure configuration parameter of the searchable unit, and training the hybrid super-network according to a sampling result and a training sample to obtain a trained hybrid super-network; the training samples correspond to label values; and searching a target sub-network meeting a preset condition from the trained hybrid super-network to serve as a compressed neural network model. According to the embodiment of the invention, the model compression precision and the model compression effect can be improved, and the model compression cost can be reduced.
Owner:SHENZHEN MICROBT ELECTRONICS TECH CO LTD

An adaptive quantization double-teacher distillation federated learning method for heterogeneous environment

PendingCN122287790APersonalizationEngineering
This invention provides an adaptive quantization dual-teacher distillation federated learning method for heterogeneous environments, belonging to the fields of adaptive quantization, model compression, knowledge distillation, and federated learning. Its technical solution includes the following steps: S1, system initialization phase; S2, user extraction and model broadcasting phase; S3, local model training phase; S4, adaptive quantization phase; S5, global model update phase; S6, termination judgment phase, when the training rounds reach their maximum value. T If the condition is met, the federated learning task is terminated; otherwise, proceed to S2. This invention is applicable to heterogeneous edge device scenarios with limited computing resources and communication bandwidth, such as smart IoT devices, enabling high-precision, low-communication personalized federated learning while ensuring privacy and security.
Owner:NANTONG UNIV

Transformer cooperative distillation incremental learning method and system based on timing consistency

The application belongs to the field of artificial intelligence model compression and edge deployment in intelligent operation and maintenance and fault diagnosis of power equipment, and discloses a transformer cooperative distillation incremental learning method and system based on time sequence consistency, which comprises the following steps: a teacher model is used to screen a no-label sample set to obtain a pseudo-label sample set, the pseudo-label sample set is combined with an original label sample set to obtain a distillation training sample set; a multi-mechanism cooperative distillation training scheme containing soft and hard label joint distillation, time sequence consistency distillation and multi-task distillation is constructed, and a student model is subjected to distillation training through the multi-mechanism cooperative distillation training scheme; and a cooperative distillation incremental learning scheme in which a cloud end continuously updates a teacher model to adapt to new multi-source monitoring data and a student model after edge distillation generates a pseudo-label to expand the distillation training sample set is constructed. The application can be used for realizing high-precision and low-cost training and continuous updating of a student model in a resource-limited terminal or an online scene.
Owner:ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID SHANDONG ELECTRIC POWER COMPANY

Multimodal large model compression method and device based on sparse codebook quantization

The application provides a multimodal large model compression method and device based on sparse codebook quantization, and relates to the technical field of computer vision, wherein the method comprises the following steps: optimizing a visual encoder to make the weight significance distribution of an arbitrary large-scale visual-language model more concentrated; evaluating the influence of each layer weight of the optimized language model on the output through second-order information, and dynamically allocating the code word quantity of each weight group according to the significance; determining the optimal sparse combination from a large-scale codebook by adopting a two-stage strategy of high-level candidate search and low-level subset refinement; and completing model quantization by combining the combination and sparse coefficients, so as to balance the compression and inference performance. The dynamic code word allocation and hierarchical search method based on sparse coding does not need additional training, can adaptively allocate the optimal sparse code word combination, can keep the model expression ability under extremely low bits, and can improve the quantization compression efficiency and inference performance.
Owner:TSINGHUA UNIVERSITY

Compression method, device and equipment for cache activation value in model fine tuning algorithm

The invention provides a compression method, device and equipment for a cache activation value in a model fine tuning algorithm, and relates to the technical field of computers, and the method comprises the steps: obtaining the cache activation value in the model fine tuning algorithm, and a compression error and an abnormal value corresponding to the cache activation value; determining a compression mode according to the compression error and the abnormal value distribution number; and compressing the cache activation value in the model fine tuning algorithm according to the compression mode to obtain a model compression result in the model fine tuning algorithm. According to the technical scheme, the compression mode for compressing the cache activation value in the model fine tuning algorithm is determined according to the compression errors and the abnormal value distribution number of different operators in the model fine tuning algorithm, and the cache activation value in the model fine tuning algorithm is compressed based on different compression modes, so that the model compression result is obtained; the model compression accuracy is improved, and the model compression efficiency is improved.
Owner:TSINGHUA UNIVERSITY