Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30results about How to "Improve reasoning efficiency" patented technology

A method, apparatus, device and storage medium for processing a request

PendingCN122088690AImprove reasoning efficiencyAvoid invalid occupationResource allocationInference methodsProcessingDatabase
This invention discloses a request processing method, apparatus, device, and storage medium, comprising: when the matching result is determined to be a matching failure, inferring the user request using a synchronous direct transmission mode through the target instance to obtain a first key-value cache, and determining the storage method of the first key-value cache based on its cost; when the matching result is determined to be a match with the target key-value cache, inferring the user request using a prefetch incremental mode through the target instance to obtain a second key-value cache, and directly saving the second key-value cache to a storage center. In the synchronous direct transmission mode, the cost of the key-value cache obtained through user request inference is calculated, and key-value caches that meet the cost requirements are saved to the storage center, avoiding invalid space occupation in the storage center and improving the inference efficiency of subsequent user requests. In the prefetch incremental mode, when a key-value cache hit meets the cost requirements, it is directly used to improve inference efficiency and avoid redundant calculations and waste of computing resources.
Owner:SHANGHAI YUNSUI TECHNOLOGY CO LTD

A method and system for identifying features in two-dimensional engineering drawings based on dynamic query-guided consistent projection.

This application discloses a method and system for identifying 2D engineering drawing elements based on dynamic query-guided consistent projection, proposing a dynamic query-guided consistent projection decoder. This decoder achieves a single forward propagation projection from a random noise vector to an accurate element state vector. The method first initializes a set of vectors following a predetermined noise distribution as initial dynamic queries. These queries are fed into the dynamic query-guided consistent projection decoder (DQ-CPD). At each layer of the decoder, the current dynamic query vector not only interacts with other queries through a self-attention mechanism, but it is also input as a controller into a "contextualized hypernetwork." This hypernetwork generates a set of dedicated weight parameters in real time based on the current inference information carried by the query vector. This method organically combines dynamic queries, dynamic context, and one-step generation theory to construct a novel, efficient, accurate, and adaptive element identification method.
Owner:BEIJING INST OF ARCHITECTURAL DESIGN

A method for predicting short-window gamma-gamma turbulence parameters in satellite-to-ground laser communication

ActiveCN121907375Bbreak through dependenceavoid lostSatellite communication transmissionTransmission monitoringEngineeringCommunications receiver
This invention discloses a short-window gamma-gamma turbulence parameter prediction method for space-to-ground laser communication, belonging to the technical field of space-to-ground laser communication and atmospheric turbulence channel parameter prediction. This method addresses the problems of traditional methods, such as strong dependence on long observation windows, significant degradation in prediction accuracy under short-window scenarios, and insufficient robustness under low signal-to-noise ratio conditions. It establishes a short-window observation model for the space-to-ground optical link and a gamma-gamma channel statistical model, constructs time-dependent short-window training data, and designs a short-window gamma-gamma network. Temporal features are extracted through a convolutional backbone, and a scintillation exponential physical regularization auxiliary head is used to achieve joint parameter prediction under physical constraints, outputting predicted gamma-gamma distributed parameters. This method can achieve high accuracy and good stability in parameter prediction under short observation windows and low signal-to-noise ratio conditions, and can be used for turbulence channel state characterization at the space-to-ground laser communication receiver.
Owner:CHANGCHUN UNIV OF SCI & TECH

Large model inference scheduling method and system, storage medium and computer program product

ActiveCN121809702BImprove stabilityload balancing
The application discloses a large model inference scheduling method and system, a storage medium and a computer program product, and relates to the technical field of large model inference scheduling. The application changes the "blind pushing" mode of the gateway in the traditional scheme, and adjusts to be dominated by the computing node. The computing node can dynamically and comprehensively perceive state data of the node, such as processor state and inference task state borne by the node. Based on the state data, the performance index of the computing node can be determined. The computing node actively reports the performance index to the gateway. The gateway can allocate target computing nodes for target inference tasks to be allocated based on the latest performance indexes of the computing nodes, better realize load balancing, and then improve the overall throughput of the system, improve inference efficiency, and improve the stability of inference service.
Owner:IFLYTEK CO LTD

Inference optimization methods, devices, and storage media for generative diffusion models

PendingCN122088667AImprove reasoning efficiencyenhance reasoning abilityProgram initiation/switchingResource allocationAlgorithmTheoretical computer science
This application provides an inference optimization method, apparatus, and storage medium for a generative diffusion model. The optimization method includes one or more of the following: first optimization of the inference parts of the U-Net and VAE in the generative diffusion model through code recompilation; second optimization of the performance bottleneck region in the VAE through data layout transformation; and multi-level operator optimization of key computational operations in the generative diffusion model based on performance bottleneck analysis. Through the above technical solutions, this invention can improve the generative diffusion model on a heterogeneous acceleration platform, aiming to locate the performance bottleneck part in the model through performance profiling, improve the inference efficiency on the heterogeneous acceleration platform through data layout transformation, and systematically improve the model's inference performance by utilizing multi-level optimization schemes from the graph level, data level, to the operator level.
Owner:DAWNING INT INFORMATION IND CO LTD

A cross-modal data processing system for safe operation of hydrogen refueling stations

PendingCN122286340AReduce labeling costshigh quality conversionData processing systemHandling system
This invention discloses a cross-modal data processing system for the safe operation of hydrogen refueling stations, including a raw data acquisition and processing module, a full-variable safety scanning module, a physical relationship coupling diagnosis module, an adaptive operating condition clustering module, a question-answer pair construction module, a hydrogen refueling station time-series command data acquisition module, and a fault type identification module. By constructing a three-layer semantic enhancement logic and model adaptation strategy, this invention effectively solves the technical problems in the prior art, such as the lack of supervision signals in the raw data of hydrogen refueling stations, the difficulty in identifying hidden faults under complex operating conditions, and the difficulty in adapting heterogeneous feature space models.
Owner:CHONGQING UNIV

Image inference method, device, apparatus and storage medium

The application provides an image reasoning method, device and equipment and a storage medium. The method comprises the following steps: inputting a user-inputted static image to be reasoned and a natural language prompt into a trained task routing module, identifying at least one task category corresponding to the natural language prompt, and activating a trained expert agent corresponding to each task category; inputting the static image to be reasoned and the natural language prompt into each trained expert agent for first reasoning, obtaining a shared semantic representation of each task category, and inputting the shared semantic representation into a collaborative memory pool for graph fusion to generate a consensus representation of the static image to be reasoned; and inputting the consensus representation of the static image to be reasoned and the natural language prompt into a trained target agent for second reasoning to obtain a final answer corresponding to each task category of the static image to be reasoned. The application avoids interference between tasks and improves the generality, reasoning efficiency, reliability and stability.
Owner:CHINA MOBILE JIUTIAN ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD +1

A cloud-edge collaborative and adaptive MoE-based video efficient analysis method

PendingCN122368866AIncrease the compression ratioefficient compressionFeature extractionEngineering
This invention discloses a high-efficiency video analysis method based on cloud-edge collaboration and adaptive MoE. The method includes: when an edge device performs a video analysis task, it first constructs a meta-request containing current task data and device resource constraints, and sends it to the cloud; the cloud uses a pre-trained basic model to extract features from the video data and clusters the data based on the feature space, dividing complex video scenes into multiple semantic subdomains; a lightweight expert model is trained in the cloud for each semantic subdomain; a routing decision model is further trained in the cloud; after deploying an expert model pool and routers on the edge device, for input data, the router first calculates the matching degree of each expert model and performs a comprehensive evaluation based on its computational cost. Through knowledge distillation, the large-scale basic model is decomposed into lightweight expert models, achieving a significant compression of the model size, enabling efficient deployment on resource-constrained edge devices.
Owner:NANJING UNIV OF POSTS & TELECOMM

Generative Recommendation System and Method Based on Quantized Vector Retrieval and Large Language Model

This invention discloses a generative recommendation system and method based on quantized vector retrieval and LLM, belonging to the technical field of computer information processing and artificial intelligence recommendation systems. The method includes: Step S1, extracting continuous collaborative features of users and items; Step S2, constructing and pre-training an AHPQ segmenter quantization module, mapping continuous collaborative features to discrete ID token sequences and optimizing the quantized codebook; Step S3, using the pre-trained codebook vector to collaboratively initialize the LLM's token embedding layer; Step S4, fine-tuning the LLM using a freeze-adapt strategy, and then generating user preference vectors through a collaborative semantic fusion module; Step S5, calculating the similarity between the user preference vector and the item database vector, and retrieving Top-K items based on cosine similarity as the recommendation result. This invention solves the problems of ID discretization difficulties and collaborative signal loss in LLM recommendation, preserving the high-order collaborative structure while leveraging the semantic reasoning capabilities of LLM, and simultaneously achieving efficient reasoning and cold-start generalization.
Owner:HARBIN INST OF TECH AT WEIHAI

A method and system for improving completion parameters for understanding user input

PendingCN122114174AAccurately capture potential needsavoid misreadingDigital data information retrievalNatural language data processingUser inputEngineering
The application relates to the technical field of industrial manufacturing, and discloses a method and system for improving the completion parameters of understanding user input, which comprises the following steps: inputting an industrial standardized file, generating a semantic label, reasoning through the industrial standardized file and the semantic label, obtaining an optimized industrial standardized file, and obtaining an industrial rule set and a text parameter set input by a user. The application realizes accurate semantic label generation based on a WordNet synonym set, solves the polysemy ambiguity problem, calculates the similarity between a candidate parameter and a user conversation through an improved Levenshtein distance algorithm and a compensation mechanism, accurately captures the potential demand of the user, improves the consistency with the actual input intention of the user, multiplies the initial weight value, the click rate and the time decay factor to sort the candidate parameters, balances the instant production demand and the historical experience in the industrial scene, and significantly improves the accuracy, adaptability, efficiency and interpretability in four dimensions.
Owner:BEIJING INFORMATION TECH BOTE INTELLIGENT TECH CO LTD

A task execution method and device, electronic equipment and storage medium

PendingCN122274957AReduce data processing complexityImprove inference speedMotion controlEmbodied intelligence
This invention provides a task execution method, apparatus, electronic device, and storage medium, relating to the field of embodied intelligence technology. The method includes: determining redundant feature components from the first task reference information based on the feature sources of feature components and the importance values ​​of each feature source, wherein the importance value of each feature source represents the magnitude of the influence of the feature component of that feature source on the successful output of the robot's expected action by the first VLA model; performing redundancy removal processing on the redundant feature components in the first task reference information to obtain processed first task reference information; inputting the processed first task reference information into the first VLA model to obtain the predicted action of the robot output by the first VLA model; and controlling the robot to execute the task based on the predicted action. Applying the solution provided by this invention can improve the robot's motion control frequency, thereby improving task execution performance.
Owner:BEIJING GALBOT AI CO LTD

An AI application development platform that integrates convolutional networks

This invention relates to the field of artificial intelligence technology and discloses an AI application development platform that integrates convolutional networks. The platform includes: a model parsing module for parsing the model and calculating the spatial importance weights and stability coefficients of the channels; a policy generation module for generating a policy table containing load levels and binary masks based on the weights and coefficients; and an application building module for encapsulating the computing power control unit to generate the target application. During program execution, the computing power control unit determines the policy level based on the hardware state and uses masks to control the data flow to either enter the neural network processor for convolution or enter the graphics processing unit for motion vector multiplexing. This invention achieves heterogeneous hardware collaboration and dynamic policy fusion through the development platform, solving the high power consumption problem of frame-by-frame full computation and improving the inference efficiency of the application and the device's battery life.
Owner:YISHU TECHNOLOGY (TIANJIN) CO LTD

Long text parallel inference method and device based on linear attention

PendingCN122242722Areduce complexityShorten the durationBiological modelsInference methods
This application relates to a method and apparatus for parallel reasoning of long texts based on linear attention. Applied to any parallel computing device in a distributed architecture, the method includes: converting a long text instruction to be processed into at least one embedded feature vector; constructing a local attention decay matrix and a global attention decay factor based on the at least one embedded feature vector and task attribute information of the long text instruction; mapping the at least one embedded feature vector to at least one core matrix based on a weight matrix; calculating the local attention matrix and the global attention matrix; constructing a total attention matrix for at least one embedded feature vector based on the local attention matrix, the global attention matrix, and linear attention; and calculating the unprocessed tokens or output text data corresponding to the long text based on task attribute information and the total attention matrix. This method can reduce the complexity of attention calculation in the long text reasoning process and shorten the computation time.
Owner:SHANGHAI XIYU JIZHI TECH CO LTD

An industrial product image angle detection and correction method based on deep learning

ActiveCN115511827BEliminate periodic abrupt changes in angle lossStrong industrial image texture feature extraction capabilityImage enhancementImage analysisPattern recognitionEngineering
This invention relates to a deep learning-based method for angle detection and correction of industrial product images, comprising the following steps: Step 1, acquiring industrial product image information; Step 2, building an industrial product image angle detection neural network model and training the network model using the industrial product image information; Step 3, after the neural network model training is completed, loading the trained model parameters, obtaining the result feature map through model forward operation, restoring it according to the labeled format to obtain the network model prediction result; Step 4, correcting the angle of the industrial product image according to the target position and angle predicted by the neural network model, and correcting the detected industrial product image angle to a uniform orientation.
Owner:BEIJING DAHENG IMAGE VISION CO LTD

An interpretable low-light image enhancement method

PendingCN122289092AImprove visual qualityphysically explainablePattern recognitionImaging processing
This invention discloses an interpretable low-light image enhancement method. It constructs an unpaired learning framework, EDC-Net, with enhancement and degradation branches forming a closed loop. The enhancement branch adaptively enhances the brightness and contrast of the low-light image through pixel-level affine mapping, while the degradation branch predicts physically meaningful degradation parameters such as color gain and exposure scaling. The enhanced image is then reconstructed into a low-light image through a chain of explicit degradation operators, forming a cyclic consistency constraint. Training employs a three-stage progressive optimization strategy: degradation warm-up, master adversarial training, and fine-tuning. During the inference phase, only the lightweight enhancement branch is run. This invention achieves physical interpretability and controllability of the enhancement process, avoids overexposure and artifacts, improves training convergence stability, and significantly increases inference efficiency. It is suitable for real-time image processing scenarios such as nighttime surveillance and autonomous driving.
Owner:JINLING INST OF TECH

Methods and electronic devices for generating house renovation drawings

This disclosure relates to a method, electronic device, medium, and program product for generating house renovation drawings. The method includes: acquiring a house boundary map and prompts describing a user's renovation needs; parsing the prompts based on a large model to determine at least one first objective to be generated and at least one second objective to be generated after the at least one first objective; generating multiple first results representing multiple generation schemes corresponding to the at least one first objective based on constraints of the house boundary map and the at least one first objective; selecting at least one second result from the multiple first results based on spatial syntax constraints; generating multiple third results representing multiple generation schemes corresponding to the at least one second objective based on constraints of the at least one second result and the at least one second objective; determining multiple first house renovation drawings representing multiple combinations of the at least one second result and the multiple third results; and determining a target house renovation drawing from the multiple first house renovation drawings.
Owner:KE COM (BEIJING) TECHNOLOGY CO LTD

A model reasoning method, apparatus, electronic device, and storage medium

PendingCN122088683AImprove reasoning efficiencyaccurate mappingInference methodsPhysical realisationAlgorithmTheoretical computer science
This application discloses a model inference method, apparatus, electronic device, and storage medium, belonging to the field of deep learning technology. The method includes: determining the inference strategy of the deep learning model based on the performance parameters of a hardware accelerator and the hierarchical performance parameters of the deep learning model; wherein the hierarchical performance parameters include: the hierarchical structure of the model layers and the inter-layer dependencies between model layers. The technical solution of this application, by comprehensively considering the performance parameters of the hardware accelerator and the hierarchical performance parameters of the deep learning model when determining the inference strategy, can achieve a precise mapping between the performance characteristics of the hardware accelerator and the computational requirements of deep learning, thereby improving the inference efficiency of the deep learning model and optimizing the inference strategy.
Owner:ANYSMART TECH CO LTD

Unmanned aerial vehicle cluster distributed model inference method, device and equipment

PendingCN122287924AGive full play to distributed computing capabilitiesquick responseUncrewed vehicleEngineering
This invention provides a method, apparatus, and device for distributed model inference in unmanned aerial vehicle (UAV) swarms, relating to the field of collaborative inference technology for UAV swarms. The method includes: acquiring environmental perception data of the UAV swarm to be inferred; obtaining a complexity assessment result based on the environmental perception data and semantic complexity evaluation; determining the inference complexity level based on the complexity assessment result and a preset inference level threshold; the inference complexity levels include low, medium, and high; determining an inference mode adapted to the environmental perception data based on the inference complexity level; low level corresponds to a lightweight inference mode, medium level corresponds to a core layer operation inference mode, and high level corresponds to a cluster splitting balanced inference mode based on matching node computing power with sub-model computing volume; and performing inference on the environmental perception data based on the inference mode to obtain the inference result. This invention can improve the inference efficiency of UAV swarm models.
Owner:STATE GRID HEBEI ELECTRIC POWER CO LTD

Neural network inference acceleration method and apparatus

PendingCN122264082Areduce processingReduce memory pressureBiological modelsInference methodsComputer hardwareEngineering
A neural network inference acceleration method and device are provided. The method includes: acquiring, by a memory having a computing unit, a video and a prompt word; acquiring, by the memory having the computing unit, at least one image related to the prompt word from the video; dividing, by the memory having the computing unit, each of the at least one image into a plurality of blocks; storing, by the memory having the computing unit, the plurality of blocks; storing, by the memory having the computing unit, position information of the plurality of blocks in each of the at least one image; acquiring, by a processor, each of the at least one image as the plurality of blocks and the position information; dividing, by the processor, each of the at least one image into a plurality of sub-blocks; acquiring, by the processor, from the memory having the computing unit, an embedded vector of a pre-stored sub-block corresponding to each of the plurality of sub-blocks, respectively; and performing an inference operation using the embedded vector and a first neural network.
Owner:SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1

Visual conversion method and device, electronic equipment, storage medium and program product

The present application relates to the technical field of computer vision, and discloses a visual conversion method and device, electronic equipment, storage medium and program product.The method comprises the following steps: acquiring multiple images; performing feature extraction on the multiple images through a visual conversion model; compressing the image features in the height dimension and the width dimension; compressing the depth features in the height dimension and the width dimension; the visual conversion model comprises an image auxiliary branch and a depth auxiliary branch in the training stage, and is removed in the inference stage; the image auxiliary branch is used for supervising the image feature compression process; the depth auxiliary branch is used for supervising the depth feature compression process; the height image features and the height depth features are fused and projected to the longitudinal area of the bird's eye view; the width image features and the width depth features are fused and projected to the transverse area of the bird's eye view; the height intermediate features and the width intermediate features are fused and compressed to obtain the bird's eye view joint features.The above scheme improves the visual conversion efficiency.
Owner:深圳魔视智能科技有限公司

A dequantization processing method and apparatus, an electronic device, and a storage medium

PendingCN122287726AImplement inverse quantization processing methodImprove the efficiency of matrix multiplication operationsAlgorithmData mining
This disclosure provides a dequantization processing method, apparatus, electronic device, and storage medium, relating to the field of artificial intelligence technology, and particularly to model inference technology. The specific scheme is as follows: The input feature matrix and model weight matrix are respectively subjected to grouped quantization processing to obtain an input feature group matrix and a model weight group matrix composed of multiple sub-matrices, and the original dequantization coefficients corresponding to each sub-matrix; A chained division transformation is performed along the common accumulation dimension on the original dequantization coefficients corresponding to all sub-matrices in the input feature group matrix to obtain first-type preprocessed dequantization coefficients; A chained division transformation is performed along the common accumulation dimension on the original dequantization coefficients corresponding to all sub-matrices in the model weight group matrix to obtain second-type preprocessed dequantization coefficients; Based on the first-type and second-type preprocessed dequantization coefficients, the matrix product result of the input feature group matrix and the model weight group matrix is ​​dequantized.
Owner:KUNWANG (SHANGHAI) TECH CO LTD

Image processing method, image processing apparatus, and computer program product

According to the image processing method, image processing apparatus, and computer program product of the present invention, an image segmentation model with adjustable boundary smoothness can be obtained, and the accuracy and efficiency of segmentation inference can be improved while improving boundary smoothness. The image processing method of the present invention utilizes a deep neural network-based image segmentation model to perform segmentation inference on a specific region contained in a medical image, comprising: a model training step, in which labeled image data and smoothness constraint parameters for constraining the boundary smoothness of the segmentation inference result are input to train the image segmentation model, thereby obtaining an image segmentation model capable of adjusting the smoothness constraint parameters; and a segmentation inference step, in which the trained image segmentation model, the input smoothness constraint parameters, and the medical image are used to perform segmentation inference on the medical image and output the segmentation inference result.
Owner:CANON MEDICAL SYST CORP

A dual-target scientific creativity generation method based on hierarchical reinforcement learning

PendingCN122174914AImprove reasoning efficiencyimprove noveltyBiological modelsTraining phaseData set
The application discloses a double-target scientific creativity generation method based on hierarchical reinforcement learning, aiming to improve the novelty and feasibility of generated content simultaneously. The method first uses a closed-source large language model to score academic review texts, builds a consensus dataset and trains novelty and feasibility reward models. In the training phase, the open-source model to be optimized is used as a strategy model. Through hierarchical sampling, novelty and feasibility-driven subgroups are constructed. The trained reward models are used to calculate the subgroup rewards and global rewards for the samples. Then, the subgroup advantages and intergroup advantages are estimated, and the final advantage value is obtained by fusion to optimize the GRPO objective function. Through hierarchical sampling and double advantage estimation mechanism, the method explicitly decouples and fuses conflicting goals, enabling the model to directly generate high-quality creativity without complex iterations during reasoning, achieving efficient and balanced scientific creativity generation.
Owner:ZHEJIANG UNIV

Knowledge graph completion method based on dynamic routing and double-channel reasoning

A knowledge graph completion method based on dynamic routing and dual-path reasoning, belonging to the field of artificial intelligence and knowledge graph technology, includes the following steps: First, extract modality-specific feature representations from the structural data, entity association visual data, and text description data of the knowledge graph, and unify the representations of each modality to the same dimension; Second, use the topological structure representation as a semantic anchor point to guide the visual and text modalities to perform asymmetric alignment towards the anchor point, obtaining multimodal entity representations in a unified semantic space; Third, evaluate the structural determinism of the query based on the multimodal entity representations to obtain the query confidence, and dynamically allocate reasoning paths to the query based on the confidence; Fourth, predict the query through two complementary reasoning paths, dynamically fuse the prediction results of each path, and output the completion result. This invention improves the reasoning reliability and computational efficiency in multimodal knowledge graph completion tasks.
Owner:CHINA JILIANG UNIV

A tongue disease risk prediction method and system based on prompt mutual learning

ActiveCN121171582Bretain knowledgeperformance maximizationFeature vectorMedicine
The application discloses a tongue disease risk prediction method and system based on prompt mutual learning. The method comprises the following steps: training a visual language teacher model based on a tongue data set through prompt learning and consistency loss, and saving a text feature vector generated by the visual language teacher model; initializing at least two student models for mutual learning, and introducing a KL divergence loss of the visual language teacher model in the mutual learning; and multiplying an image feature vector of a tongue image to be predicted extracted by any one student model with the text feature vector generated by the visual language teacher model to output a disease risk prediction result. The application effectively improves the accuracy and reliability of tongue disease risk prediction by fusing prompt learning and mutual learning distillation technology.
Owner:SOUTH CHINA UNIV OF TECH

Http / 3 encrypted traffic intelligent identification method based on empty load removal packet cluster representation unit

PendingCN122293379AAdapt to downstream classification taskslow costPlaintextData pack
This invention proposes an intelligent identification method for HTTP / 3 encrypted traffic based on empty payload packet cluster representation units. The specific steps are as follows: 1) Preprocess the original HTTP / 3 encrypted traffic, remove empty payload data packets, and mask plaintext fields related to the environment such as IP address, port number, and certificate fingerprint to filter noise; 2) Construct packet cluster representation units for the traffic based on bidirectional packet analysis, divide the cleaned traffic into multiple packet clusters according to a fixed time window, extract the uplink and downlink packet subclusters in each cluster, calculate the statistical characteristics of the subclusters, and extract the header and end ciphertext fragments of the packets to form a bidirectional packet cluster structure. Convert the packet cluster structure into a token sequence using the BPE method; 3) Use the BERT model for pre-training and subsequent model tuning. During pre-training, set up mask prediction tasks and homogeneous packet cluster prediction tasks. During fine-tuning, use the output fields of the CLS block in the pre-trained model as the classification basis for supervised training and output the tuned model.
Owner:SOUTHEAST UNIV

Quantization methods, devices, equipment, media, and programs for visual language models

PendingCN122088707ASolve the problem of low quantization accuracyHigh quantitative accuracyBiological modelsInference methodsLinguistic modelAlgorithm
This invention discloses a method, apparatus, device, medium, and program for quantizing a visual language model. The method includes: quantizing the visual modules of the visual language model during an offline quantization stage; compensating for quantization errors in the visual modules using a quantization compensator of the visual language model; and performing staged quantization processing on the language modules of the visual language model. The technical solution of this invention can improve the quantization accuracy of large-scale visual language models, thereby improving the inference efficiency of large-scale visual language models and reducing the inference time and computational cost of large-scale visual language models.
Owner:SHANGHAI SUIYUAN TECH CO LTD

An automatic driving, model fine-tuning method, device, system and vehicle

ActiveCN121404318BImprove cross-scenario reasoning flexibilityImprove reasoning efficiencyInference methodsSimulationMultimodal data
This application provides an autonomous driving method, apparatus, system, and vehicle for model fine-tuning, relating to the field of autonomous driving. The method includes: acquiring system prompts and multimodal data of an autonomous driving task, where the system prompts indicate the selection of a reasoning mode based on scene information; processing the multimodal data and system prompts using a Virtual Logic Analyzer (VLA) model to obtain current scene information of the autonomous driving task and the corresponding current reasoning mode; processing the multimodal data under the current reasoning mode of the VLA model to obtain a current processing result; and executing the autonomous driving task according to the current processing result. This solution addresses the problems of insufficient cross-scene reasoning flexibility and low reasoning efficiency of the VLA model.
Owner:NEW ZIGUANG GROUP CO LTD