Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

476 results about "Model compression" patented technology

Road safety early warning method and system based on mixed precision quantification visual large model

The invention discloses a road safety early warning method and system based on a mixed precision quantification visual large model, and the method comprises the steps: collecting road traffic safety videos and pictures, and carrying out the preprocessing, data enhancement and marking, thereby forming a diversified data set; a pre-trained visual large model is selected as a teacher model, after fine tuning, output layer and middle layer knowledge is extracted, key features are weighted, and meanwhile, a lightweight neural network is taken as a student model, same input is received, and prediction and middle feature maps are output. And inputting data into the two models and the student model to carry out mixing precision quantification forward propagation, constructing a total loss function containing tasks, knowledge distillation and quantification learning loss, and updating parameters through back propagation. And after training is completed, exporting a quantitative model, and deploying the quantitative model to an edge computing platform to realize safety early warning. The lightweight model can realize rapid reasoning on edge equipment such as a vehicle-mounted road side, and the problem that performance and efficiency are difficult to consider in a traditional model compression method is solved.
Owner:HARBIN INST OF TECH

Intelligent security remote inspection and control method based on AI

The invention relates to the technical field of security management, in particular to an AI-based intelligent security remote inspection and control method, which comprises the steps of front-end intelligent equipment deployment, intelligent inspection execution, data processing and intelligent analysis and remote control and linkage response. Compared with the defects of incomplete information and high false alarm rate due to the fact that equipment appearance anomaly detection mainly depends on single-modal data in the prior art, the scheme adopts multi-modal data preprocessing and feature extraction, constructs a cross-modal alignment network architecture, dynamically fuses laser radar, RGB-D images and IMU data by using a self-attention mechanism, and improves the detection accuracy of the equipment appearance anomaly. According to the method, more accurate equipment appearance abnormity identification is realized by combining environment self-adaptive adjustment, and the model is compressed and deployed to an edge end through a knowledge distillation technology, so that real-time reasoning and analysis are realized, the intelligent level of security inspection and the accuracy of abnormity detection are remarkably improved, the false alarm rate is reduced, and the practicability and reliability of the system are enhanced.
Owner:GUANGZHOU YOUSEN INFORMATION TECHNOLOGY CO LTD

End-side cloud cooperation system based on hybrid scheduling strategy

The invention relates to the technical field of artificial intelligence and edge computing, in particular to an end-side cloud collaboration system based on a hybrid scheduling strategy. The end-side cloud cooperation system based on the hybrid scheduling strategy comprises a model segmentation module, a state sensing module, a reasoning scheduling module, a model compression and deployment module, an edge optimization engine and a communication synchronization module. According to the end-side cloud cooperation system based on the hybrid scheduling strategy, elastic deployment and dynamic reasoning of a complex model in a multilayer heterogeneous environment are supported, the system state can be sensed in real time, a scheduling path can be optimized, and the robustness of the system is improved; and in combination with model compression and edge operator optimization, the reasoning efficiency is remarkably improved, the communication load is reduced, and then low-delay cooperation between the terminal and the cloud is ensured.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Digital twin drive multi-source heterogeneous data fusion and intelligent analysis method for oil and gas pipeline construction

The invention discloses a digital twin-driven multi-source heterogeneous data fusion and intelligent analysis method for oil and gas pipeline construction. The method comprises the following steps: collecting multi-source heterogeneous data related to oil and gas pipeline construction; fusing data by using a space-time diagram neural network and constructing a dynamic offset compensation model to correct coordinate drift; constructing a data credibility dynamic evaluation chain, performing cross validation on sensor data to generate a confidence coefficient weight, and complementing missing data through a GAN complementing module to form a credibility enhanced data set; designing a layered distillation compression algorithm to compress the PINN model into a TinyLSTM engine, and developing a model increment updating protocol to realize efficient collaboration of a cloud end and an edge end; and training and optimizing a TinyLSTM engine by using the data set with enhanced credibility. According to the method, the defects in the aspects of data dynamic correction, credibility evaluation and complementation, model edge deployment real-time response and the like in the prior art are overcome, and the method has remarkable innovativeness and practicability.
Owner:INNER MONGOLIA WESTERN NATURAL GAS CO LTD +1

Image big data classification and identification method and system based on deep learning

The invention relates to the field of computer vision and deep learning, and discloses an image big data classification and recognition method and system based on deep learning, and the method comprises the steps: generating a gating matrix through the extraction of an image frequency domain energy coefficient, compressing a convolution kernel through the combination of asymmetric tensor decomposition, and carrying out the self-adaptive training through cross-modal semantic alignment and a meta-learning task. Efficient classification reasoning of dynamic path selection is realized, and the precision and the calculation efficiency are improved; the system comprises a frequency domain analysis module, a dynamic sparse gating module, an asymmetric tensor decomposition module, a meta-learning task generation module, a cross-modal alignment module and a dynamic inference engine module. According to the method, through cross-modal semantic alignment and meta-learning task optimization, in combination with lightweight parameter storage and edge calculation path selection, fine-grained classification precision improvement, model compression and high-efficiency reasoning are realized, and the calculation efficiency and generalization ability in a complex scene are remarkably enhanced.
Owner:BEIJING NANSHAN TONGXING TECHNOLOGY CO LTD

Large model compression method and device, task processing method and equipment and storage medium

The invention relates to the technical field of model compression, and provides a large model compression method and device, a task processing method and equipment and a storage medium, and the large model compression method comprises the steps: carrying out the layer-by-layer quantification of a linear layer of a to-be-compressed initial large model, and obtaining a first large model; the initial large model is a pre-trained large language model constructed based on an expert hybrid architecture; performing route calibration on each expert sub-model in the first large model to obtain a second large model; in the reasoning process of the second large model, based on the task type of a to-be-executed target task, the importance of each expert sub-model in the task type is evaluated; and performing dynamic pruning on each expert sub-model based on importance so as to compress the second big model. Through a compression mode of combining static quantification and dynamic pruning, on the basis of ensuring the model performance, the memory and calculation overhead required by large model reasoning can be reduced, and efficient operation of the large model on light-weight equipment with limited video memory resources is facilitated.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

PDF drawing identification and information structured extraction method

The invention discloses a PDF (Portable Document Format) drawing recognition and information structured extraction method. The method comprises the following steps: generating a high-resolution bitmap through image preprocessing; positioning and classifying a text region, a table region and a symbol region in the drawing based on a target detection model of transfer learning; hough transform is combined with SIFT feature matching to identify engineering symbols, and sub-pixel positioning is realized through an RANSAC algorithm; after the oblique text is corrected through affine transformation, the content is extracted through OCR; reconstructing a table structure based on OPTICS clustering and projection analysis; constructing an RDF knowledge graph according to a coordinate association rule; and using U-Net difference to detect and position an omission area and complementing the omission area. According to the method, deep learning and image processing technologies are fused, the problems of low rotating text recognition rate, table structure loss and semantic association deficiency in a traditional method are solved, through lightweight model compression and TensorRT acceleration, the analysis accuracy is remarkably superior to that of the traditional method, and the method can be widely applied to the fields of constructional engineering, petrochemical engineering and the like and has wide application prospects. And the drawing information processing efficiency and the data integrity are improved.
Owner:ZHEJIANG THERMAL POWER CONSTR CO LTD

Three-dimensional model compression transmission method and system based on dynamic feature perception

The invention relates to a three-dimensional model compression transmission method and system based on dynamic feature perception, and relates to the technical field of three-dimensional live-action modeling, and the method comprises the steps: obtaining the original three-dimensional data of a three-dimensional model, carrying out the multi-modal feature extraction based on a model scene, dividing a key region, a transition region and a non-key region, and generating a three-dimensional model partition map, compressing each region according to the partition map label and a preset compression algorithm to generate a multi-resolution LOD sequence, determining a transmission data hierarchy in combination with a network state and equipment parameters, and finally rendering the model according to the transmission data hierarchy, user behavior information and an environment state to obtain a target three-dimensional model. The technical effects of effectively compressing, transmitting and rendering the three-dimensional model according to various factors such as different area characteristics, network states, equipment parameters and user behaviors of the model and improving the transmission efficiency and the rendering quality are achieved.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Retraining-free pruning and recombination method and system for sparse expert hybrid large model

The invention discloses a retraining-free pruning and recombination method for a sparse expert hybrid large model, and belongs to the technical field of large model compression and optimization. The method aims at solving the problems that due to the fact that an existing sparse expert hybrid (SMoE) model needs to load all expert parameters, memory occupation is too high, and deployment is difficult. According to the method, firstly, redundant experts are identified and pruned based on routing activation statistics; then, decomposing the pruned experts into neuron-level functional fragments, and redistributing the fragments to the reserved experts according to structural similarity; and finally, original fragments and newly distributed fragments are merged in the reserved experts through a weighted clustering algorithm, so that compact experts with fewer parameters and stronger expression ability are reconstructed. According to the method, fine-grained operation is carried out at the neuron level, the inherent representation conflict and dislocation problems among experts are effectively solved, the performance of the compressed model is remarkably improved, and reliable technical support is provided for deploying a large-scale SMoE model.
Owner:ZHEJIANG UNIV

Main control chip task scheduling and dynamic performance optimization method based on neural network

The invention relates to the technical field of chip scheduling and optimization, in particular to a neural network-based main control chip task scheduling and dynamic performance optimization method, which comprises the steps of data acquisition and feature engineering, neural network model design, simulation environment training, model compression and deployment preparation, real-time state monitoring, dynamic decision reasoning, scheduling strategy execution and performance optimization. The data acquisition and feature engineering comprises the following steps: S1, hardware index acquisition; analyzing task attributes (calculation-intensive / IO-intensive), a dependency relationship (DAG), deadline (Deadline) and a resource demand (CPU / GPU occupancy rate); collecting data during chip operation through a performance counter (IPC, cache hit rate and branch prediction error rate), a temperature sensor and a power consumption monitoring unit (PMU); the neural network scheduler can achieve the energy efficiency ratio which is 20%-40% higher than that of a traditional method (such as a CFS scheduler), meanwhile, the neural network scheduler adapts to sudden load changes, and the practicability and the application range of a main control chip are wider.
Owner:HUNAN SHENGYUN PHOTOELECTRIC TECH CO LTD

Large model dynamic compression optimization method and system based on sparse pruning

The invention relates to the technical field of large model algorithms, in particular to a large model dynamic compression optimization method and system based on sparse pruning, and the method comprises the steps: capturing original weight fluctuation data generated by resource fluctuation in reasoning, and obtaining sparse weight reference data through sparse processing; analyzing calculation complexity through model reasoning delay data, and separating reasoning delay amount caused by a model scale; dynamically controlling the model compression ratio within a preset performance range based on the delay amount and the sparse reference data, and collecting reasoning precision distribution data under different compression parameters; evaluating a model performance state under each parameter by means of a neural network simulation method, and generating a performance state simulation result; determining a model quality optimization compensation parameter based on a simulation result by combining resource fluctuation data acquired in real time in a compression process; and finally, the compression strategy is adaptively regulated and controlled through the compensation parameters, and collaborative optimization of model calculation complexity, reasoning precision and delay during dynamic change of hardware resources is realized.
Owner:NOVNET COMPUTING SYST TECH CO LTD

Equipment abnormal voiceprint detection system

The invention provides an equipment abnormal voiceprint detection system. The equipment abnormal voiceprint detection system comprises a voiceprint collection module used for collecting original voiceprint signals in real time and preprocessing the original voiceprint signals; the edge end detection module is deployed on edge computing equipment and is used for carrying out real-time anomaly detection on the voiceprint features by utilizing a one-dimensional lightweight neural network model; the abnormity credibility evaluation module is used for converting a real-time abnormity detection result into a probabilistic abnormity credibility score; the incremental data screening module is used for screening high-value samples from the real-time voiceprint data based on a dynamic density trend sensing algorithm and caching the high-value samples in the edge; the cloud model evolution module is used for carrying out evolution training on the detection model by utilizing a continuous learning mechanism; the generation and playback module is used for jointly generating pseudo samples through a variational auto-encoder VAE and a generative adversarial network GAN; and the model updating module is used for compressing the evolved cloud model and then issuing and replacing the original model in the edge end detection module so as to form a cloud-edge collaborative sustainable evolution closed loop.
Owner:ZHONGZHENG EVALUATION (SHENYANG) TECHNOLOGY CO LTD

Control system and method for intelligent engine cooling water pump actuator

The invention discloses a control system and method for an intelligent engine cooling water pump actuator, and the system integrates a multi-mode perception technology, dynamic thermal entropy feature extraction, a hierarchical reinforcement learning strategy, graph neural network collaborative optimization, adaptive model prediction control, and an edge calculation deployment and fault self-repairing mechanism. A whole-process intelligent cooling control framework is constructed; collecting multi-source sensor data, calculating dynamic thermal entropy through a hybrid mechanism, and performing weighted fusion and anomaly suppression on the data; a hierarchical control strategy is adopted to make a cooling water pump target rotating speed decision, a finite element analysis method is introduced to dynamically predict thermal stress, the thermal stress serves as a penalty term to be incorporated into a strategy optimization process, a graph neural network is constructed, and inter-subsystem coupling relation modeling and cooperative control instruction generation are achieved; a self-adaptive MPC optimization model is constructed, and a lightweight control model compression method based on knowledge distillation is adopted to adapt to vehicle-mounted edge computing resource limitation; and designing an online thermal entropy monitoring mechanism and an incremental learning self-repairing strategy.
Owner:DAFENG HAINA MACHINERY

Target lightweight detection method and system based on attention feature enhancement

The invention relates to the technical field of image target detection, and provides a target lightweight detection method and system based on attention feature enhancement, and the method comprises the steps: collecting visible light target image data; constructing a target lightweight detection model: designing a lightweight backbone network and a special attention mechanism; designing an adaptive feature fusion module; designing a detection head; designing a loss function; designing a model lightweight strategy; and training the target lightweight detection model through the training set, inputting the test set into the trained target lightweight detection model after training is completed, and outputting a target lightweight detection result. According to the scheme of the invention, the method focuses on the design of an efficient lightweight network architecture, and through the introduction of a special attention mechanism, dynamic feature fusion and a model compression strategy, the detection precision is ensured, meanwhile, the demand for computing resources is remarkably reduced, and the real-time detection of a visible light image target is realized.
Owner:NAVAL AVIATION UNIV

Home semantic segmentation model compression method based on multi-layer feature alignment knowledge distillation

The invention relates to the technical field of computer visual recognition, in particular to a home semantic segmentation model compression method based on multi-layer feature alignment knowledge distillation, which specifically comprises the following steps: establishing a teacher model taking SegformerB4 as a framework and a student model taking DeepLabV3 as a framework, and respectively extracting global context features and local detail features; enabling the student model to obtain the deep global heterogeneous features of the teacher model through a re-blocking cross attention feature alignment module, and making up the local receptive field defect of the convolutional network; a heterogeneous scale feature alignment module is adopted to reconstruct multi-scale feature maps output by different networks, and the problem of resolution mismatching caused by structural difference is eliminated; through a dual feature alignment mechanism, the student model fully learns knowledge representation of the teacher model, and the segmentation precision is improved while the size of the student model is kept unchanged. Through the process, the semantic knowledge of the teacher model is migrated to the student model, so that a lightweight but high-performance segmentation network is trained.
Owner:SHANDONG UNIV +3

Method and system for synchronously monitoring multiple physiological parameters and wearable information acquisition equipment

The invention provides a multi-physiological-parameter synchronous monitoring method and system and wearable information collection equipment. The method comprises the steps that multiple physiological signals synchronously collected through signal collection equipment are obtained and preprocessed; the preprocessed physiological signals are stored in a segmented mode according to a time sequence; each segmented physiological signal is loaded to a corresponding lightweight model for synchronous monitoring, wherein the lightweight model comprises an initial convolution module, a depth separable residual module, an output feature fusion module and a convolution attention module; the lightweight model carries out model compression through a double-teacher-guided progressive pruning distillation framework and a mixed quantification strategy, and edge side collaborative acceleration is realized in combination with a model pre-compiling technology and a multi-model segmentation scheduling strategy. According to the scheme, on the basis of ensuring the monitoring accuracy, the model complexity and the system energy consumption are greatly reduced, efficient, synchronous and real-time monitoring of multiple physiological parameters is achieved, and the method is suitable for monitoring scenes of multiple edge physiological parameters.
Owner:FUDAN UNIVERSITY

Front-end resource dynamic preloading method, system, equipment and medium

The invention provides a front-end resource dynamic preloading method, system and device and a medium, and belongs to the technical field of front-end engineering. The method comprises the following steps: constructing and training a user behavior prediction model based on a Transform time sequence model architecture, and deploying the user behavior prediction model to a browser side after compression processing; the method comprises the following steps: acquiring behavior data of a user by monitoring a user interaction event, constructing a time sequence operation sequence, and collecting network information and equipment performance information; inputting the time sequence operation sequence into the model, outputting probability distribution of the next operation, and determining related resources; based on the probability distribution, the resource type and the network information, calculating the priority of the resource, and generating a resource priority queue; according to the network information and the equipment performance information, determining a pre-loaded resource type, and adjusting a resource priority queue to generate a pre-loading strategy; a cache rule is defined, a cache configuration file is generated, and a CDN cache is used for preloading resources; and executing the preloading strategy, and carrying out rollback processing after loading fails.
Owner:SHANDONG INSPUR CLOUD GOVERNMENT INFORMATION TECHNOLOGY CO LTD

Hot line platform label model training method based on deep neural network

The invention relates to the technical field of model training, in particular to a hotline platform label model training method based on a deep neural network. The method comprises the following steps: preprocessing an input hotline work order text, and extracting channel, region, time and other structured metadata; performing semantic coding on the text by adopting a pre-trained BERT model, and dynamically fusing structured metadata and text features through a gating attention mechanism; performing field adaptive pre-training by using a hierarchical encoder, and outputting an enhanced semantic representation vector; constructing a label co-occurrence probability matrix and inputting the label co-occurrence probability matrix into the graph neural network for multi-label collaborative prediction; carrying out adaptive optimization by combining a focus loss function and a meta-learning framework; the teacher model is compressed into a student model through knowledge distillation, and the student model is deployed to an NPU hardware platform, so that real-time reasoning and online feedback closed loop are realized. According to the method, the work order classification accuracy can be remarkably improved, the order sending error rate and the processing delay are reduced, and the hotline service intelligent level and the operation and maintenance efficiency are improved.
Owner:XIAN KAILINSTE NETWORK TECHNOLOGY CO LTD

Gear machining multi-source thermal error compensation system based on tensor compression edge deployment

The invention discloses a gear machining multi-source thermal error compensation system based on tensor compression edge deployment, which fuses physically guided multi-source error modeling, tensor-based model compression, edge deployment and bandwidth sensing signal scheduling to realize low-delay and high-precision error compensation. The method comprises the following steps: firstly, establishing a three-layer system architecture comprising a cloud layer, an edge layer and a sensing layer, supporting distributed model training, edge reasoning and error compensation so as to support closed-loop operation for realizing adaptive scheduling with extremely low delay, and mainly comprising: (1) a multi-source error model, thermal and geometric errors are fused into coupling high-order representation for capturing space-time interaction through an analytical method; and (2) establishing a sensitivity analysis model based on a physical guidance tensor mode decomposition method (PG-TMDM), identifying a key error source through a three-order tensor structure, and calculating a physical perception contribution rate to retain physical interpretability. And in order to realize real-time deployment, deploying to the edge layer through model compression, model packaging and an edge deployment strategy.
Owner:CHONGQING UNIV

Traffic accident intelligent detection system and method based on YOLOv12 improved architecture

The invention relates to a traffic accident intelligent detection system and method based on a YOLOv12 improved architecture. The system comprises a YOLOv12 enhanced feature extraction network, a multi-scale detection head, a time sequence information fusion module, a real-time reasoning optimization engine and an intelligent decision fusion system, and realizes collaborative optimization of local feature enhancement and global context modeling by constructing six core technology modules and adopting collaborative learning of a C2f-Attention mechanism and deformable convolution. According to the method, a composite loss function special for traffic accidents is innovatively designed, and adaptive fusion of multi-scale features and difficult sample mining are realized through a multi-objective optimization mechanism of Enhanced Focus Loss, IoU-aware Loss and Severage-aware Loss. According to the method, the problems of low detection precision and false alarm and missing alarm caused by illumination variation, shielding and motion blur in a traffic monitoring scene are effectively solved, in the test of an AccidentsDesection YOLOv8 data set, the mAP at 0.5 reaches 91.27% and is improved by 8.6% compared with that of YOLOv8, the reasoning speed reaches 67 FPS, experimental results show that the system has excellent performance in the aspects of detection precision, real-time performance and model compression, and the method is suitable for popularization and application. The method achieves a remarkable effect in traffic accident intelligent identification, and has a remarkable technical effect and industrial application value.
Owner:JIANGSU OCEAN UNIV +1

Model quantification method and system based on stability scoring

The invention discloses a model quantification method and system based on stability scoring, and the method comprises the steps: calculating the output features of each layer of an original model and the probability distribution of the output features of each layer of a quantification model, calculating the stability value of each layer of the quantification model according to the probability distribution, and evaluating the stability of each layer of the model through the stability value. Bit width is distributed according to characteristics of each layer, the method does not depend on labels, is only based on network forward activation information, and simultaneously measures output distribution difference and direction consistency before and after quantization, and a hierarchical stability measurement index can be constructed according to a stability value to accurately identify a steady-state layer and a high-sensitivity layer for guiding precision distribution. The bit width is adaptively reduced by adopting an iterative search strategy on the basis of ascending sorting of the stability value scores; on the premise that precision and resource constraints are met, the model compression ratio is increased to the maximum extent; retraining is not needed in the whole process, and large-scale network one-key quantitative deployment is supported.
Owner:XIAN TECH UNIV

Differential privacy protection image recognition neural network training method and system

The invention provides a differential privacy protection-based image recognition neural network training method, which comprises the following steps of: applying a special random quantization scheme controlled by learnable privacy parameters to a model weight of a neural network, and constructing a total loss function containing task loss and privacy loss; and jointly optimizing the model weight and the privacy parameters in an end-to-end training mode. According to the method, the differential privacy is realized by using the endogenous randomness of the quantization process, and extra noise does not need to be injected, so that provable privacy protection is provided, and meanwhile, remarkable reduction of model precision caused by a traditional gradient perturbation method is avoided. In addition, end-to-end optimization is carried out by taking privacy parameters as learnable variables, and an optimal balance point of model utility and privacy protection is automatically sought; and through direct disturbance on the model weight, the problem of accumulation of privacy budget in the training process is avoided, efficient fine adjustment on the pre-training model is supported, and model compression and privacy protection are cooperatively realized.
Owner:SHANGHAI JIAOTONG UNIV

Large model edge deployment method, device, equipment, medium and product

The invention discloses a large model edge deployment method and device, equipment, a medium and a product, and relates to the technical field of edge computing. The method comprises the following steps: acquiring a teacher model and a student model based on a time sequence, and initializing parameters of the student model; the teacher model comprises L teacher hierarchies preselected based on an edge computing task, the student model comprises student hierarchies in one-to-one correspondence with the teacher hierarchies, and time sequence attention hierarchies located before / after the preselected hierarchies are configured in the teacher and student models; obtaining time sequence data and dividing the time sequence data to obtain training samples of different scales; obtaining a current resource state of the edge equipment, determining weights of hierarchical distillation loss and cross-scale time sequence characteristic loss according to the current resource state, and determining a time sequence attention loss weight; in the process of training the student model, performing multi-level knowledge distillation according to the level distillation loss weight, and performing time sequence alignment distillation according to the cross-scale time sequence feature loss weight and the time sequence attention loss weight; and determining a trained model according to loss of task execution of the multi-level knowledge distillation, the time sequence alignment distillation and the student model, and deploying the trained model in edge equipment after pruning. According to the method, when the model is deployed at the edge end, model compression can be realized, and dual requirements of the edge side on reasoning efficiency and reasoning precision are considered.
Owner:广东知业科技有限公司

Multi-modal model compression and distillation method and system based on causal reasoning

The invention relates to the field of multi-modal neural network model compression, and particularly discloses a multi-modal model compression and distillation method and system based on causal reasoning, and the method comprises the steps: constructing a comprehensive causal discovery module, identifying a causal dependency relationship among the multi-modal features through information theory measurement, Granger causal analysis and intervention-based verification; executing an adaptive compression engine, and performing pruning, mixing precision quantification and low-rank decomposition based on a causal relationship; a cross-modal distiller is applied, and multiple loss function combinations are adopted to maintain the relationship between modals; and implementing a dynamic optimizer to carry out hardware perception and context-sensitive reasoning optimization. According to the method, the compression decision is guided through causal reasoning, the high compression rate is achieved while the key causal path is kept, and the deployment problem of the multi-modal model in the resource-constrained environment is effectively solved.
Owner:SHENZHEN UNIV

Dynamic scheduling lightweight method and system for speculative decoding of embedded heterogeneous system

The invention discloses an embedded heterogeneous system speculative decoding dynamic scheduling lightweight method and system, and the method comprises the steps: distributing speculative decoding tasks to three processors of a CPU, a GPU and an NPU, carrying out the model compression processing of a draft model and a target model in the speculative decoding process, executing the correspondingly distributed speculative decoding tasks by each processor, and carrying out the model compression processing of the target model; a load evaluation model is constructed to evaluate the load condition of a processor in the speculative decoding process, the draft generation length is dynamically adjusted according to an evaluation result, a multi-draft parallel processing mechanism is introduced, and resources are dynamically allocated to draft generation threads. According to the method, joint optimization is carried out on model compression and calculation scheduling strategies in the speculative decoding process, the inference performance is improved, power consumption is reduced, and help is provided for solving the inference efficiency bottleneck problem of a large language model under the condition that embedded device resources are limited.
Owner:SOUTHWEST JIAOTONG UNIV

Intelligent tracking analysis system for business

The invention discloses an intelligent tracking analysis system for business, and particularly relates to the field of data intelligent analysis and business process modeling, the system accesses an enterprise internal multi-source heterogeneous system, collects structured and unstructured data related to the process, and constructs a behavior map model; and compressing and extracting key nodes and behavior fragments in the multi-role cooperation path. Simulating a future process trend by using a shadow business body, predicting a possible path and an accompanying risk thereof, and introducing a multi-agent strategy game mechanism to realize dynamic game and consensus formation of each role in path selection; and correcting the deviation behavior in the process by combining real-time path monitoring and an abnormal intervention mechanism. According to the method, dynamic modeling, prediction and intervention of a complex process can be realized, and the problems of approval lag, path conflict and the like are effectively avoided. Compared with a traditional control mode, the method has higher adaptability, flexibility and evolutionary ability, and is suitable for various high-collaboration scenes such as finance, government affairs and engineering.
Owner:NANTONG LEIHUAN TECHNOLOGY CO LTD

Method for detecting health state of household energy storage lithium battery

The invention relates to the technical field of energy storage lithium batteries, in particular to a method for detecting the health state of a household energy storage lithium battery. The method for detecting the state of health of the household energy storage lithium battery is characterized by comprising the following specific steps: S1, preprocessing charge-discharge cycle data; s2, performing multi-domain feature extraction; s3, an SA-XGBoost model oriented to static data and a TSO-LSTM-AM model oriented to dynamic data are constructed respectively, and SOH estimation is carried out on the SA-XGBoost model and the TSO-LSTM-AM model; and S4, finally, performing compression optimization on the model through knowledge distillation and quantification technologies to meet a lightweight deployment requirement. Compared with the prior art, in combination with historical charging and discharging parameters such as battery voltage, current and temperature, feature vectors are constructed through adaptive data cleaning and multi-domain feature extraction, the SOH estimation capability of static and dynamic data is optimized by adopting machine learning and a deep learning model respectively, and the deployment requirement of practical application is met by utilizing a model compression technology.
Owner:SHANGHAI PYTES ENERGY CO LTD

Private generation type large model optimization system based on hybrid genetic algorithm

The invention discloses a private generation type large model optimization system based on a hybrid genetic algorithm, belongs to the technical field of artificial intelligence, and aims to solve the technical problems of how to compress the model scale, simplify the feature dimension to improve the reasoning efficiency and maintain the model precision as much as possible. Comprising a multi-source heterogeneous data fine adjustment module, a large model construction module, a hybrid genetic algorithm optimization module, a diffusion type encoder feature extraction module, a course type noise robustness improvement module, an activation function optimization module and a feature dimension reduction and low-rank compression module, the large model is gradually and deeply optimized from the aspects of data, structure, training, robustness, activation function and model compression, and a generative large model which can fully learn financial data of the new energy vehicle enterprise and is excellent in performance, reliable in robustness and friendly to resources is produced.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Federal learning model compression method based on data-free knowledge distillation and multi-teacher knowledge distillation

The invention discloses a federated learning model compression method based on data-free knowledge distillation and multi-teacher knowledge distillation, and belongs to the technical field of federated learning. According to the technical scheme, a central server is used for pre-training a model to constrain random noise and labels, a loss function is optimized after the model is input into an image generator, and intermediate features are extracted and stored in a feature pool; and after the client uploads the model, sampling from the feature pool to generate a diversified composite image as training data, and training a local aggregation model by combining a pre-training model and a client integrated model as teachers. The beneficial effects are that the FedPET method provided by the invention combines data-free knowledge distillation and integrated knowledge distillation, and innovatively introduces channel feature exchange and multi-teacher model guidance, so that the model compression efficiency, generation diversity, data privacy protection, communication efficiency and accuracy of federal learning are significantly improved; the method has obvious advantages especially in heterogeneous data distribution and resource limited scenes, and an efficient, safe and accurate solution is provided for practical application of federated learning.
Owner:DALIAN UNIV OF TECH

Aircraft airfoil pressure self-adaptive calibration method for wind tunnel environment

The invention discloses an aircraft airfoil pressure self-adaptive calibration method for a wind tunnel environment, belongs to the technical field of pressure sensors, and aims to solve the problem that the sensitivity of the sensors is reduced during aircraft airfoil pressure testing. Comprising the steps that a flexible pressure sensor array is installed in a matched mode according to the airfoil shape of a to-be-tested aircraft, the to-be-tested aircraft is placed in a wind tunnel, wind tunnel parameters are adjusted, and pressure signals in a multi-scene environment are generated; each pressure sensor unit is regarded as an independent agent, and a calibration model is constructed by using an MADDPG algorithm; an Actor-Critic network architecture of multi-agent cooperative training is designed and obtained; and performing model compression on the trained Actor network, deploying the Actor network to an embedded system, fitting a nonlinear mapping relation between output of a pressure sensor unit and a real pressure value in real time, correcting a nonlinear error, and performing dynamic compensation on zero drift. The method is used for aircraft airfoil pressure distribution testing.
Owner:NO 49 INST CHINESE ELECTRONICS SCI & TECH GRP +1