Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

13186 results about "Training methods" patented technology

Vertical large language model training method and system in carbon neutralization field

The invention discloses a vertical large language model training method and system in the carbon neutralization field, and the method comprises the following steps: collecting data of the carbon neutralization field, carrying out the data preprocessing, constructing a carbon neutralization field knowledge base, and updating the carbon neutralization field knowledge base through a dynamic updating mechanism; performing dynamic semantic partitioning and vectorization coding on the text of the carbon neutralization domain knowledge base, and storing the text into a vector database; performing staged fine tuning on the pre-trained large language model based on a low-rank adaptation technology, wherein the fine tuning comprises general instruction fine tuning and carbon neutralization field professional fine tuning; a retrieval enhancement generation mechanism is adopted, knowledge fragments related to user query are retrieved through a vector database, and a large language model is input to generate answers. Compared with the prior art, the method has the advantages that the answer reliability is improved through conflict detection and source tracing, so that the large language model can more accurately adapt to knowledge requirements in the carbon neutralization field.
Owner:SUN YAT SEN UNIV

Target multi-modal model system and construction method, video processing model training method, and video processing method

Embodiments of the present invention provide a target multi-modal model system and construction method, a video processing model training method, and a video processing method. The video processing model training method comprises: inputting a video sample and each initial text sample into a video processing model, wherein the initial text sample is a text for performing category description on video content of the video sample; using the video processing model to perform feature extraction on the video sample to obtain a temporal motion feature and a fused image feature; using the video processing model to perform feature extraction on the initial text sample to obtain a dynamic text feature and a fused text feature; and training the video processing model on the basis of the temporal motion feature and the dynamic text feature, and the fused image feature and the fused text feature.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Polarization three-dimensional reconstruction method and system based on prior guide diffusion model

The invention provides a polarization three-dimensional reconstruction method and a polarization three-dimensional reconstruction system based on a prior guide diffusion model, which apply a diffusion model in the field of polarization three-dimensional reconstruction and improve the recovery capability of complex details and the robustness of noise interference resistance. According to the method, a two-stage training mode is adopted, and the generation quality and the calculation efficiency are balanced through step-by-step optimization. In the first stage, a VQGAN codec is independently trained to learn high-efficiency low-dimensional potential representation of an image, and direct high-cost calculation in a pixel space is avoided; in the second stage, learnable parameters of the VQGAN are frozen, a diffusion model is trained on a trained potential space, gradual denoising is guided through a priori condition, and potential features are generated and mapped back to an image space. The diffusion model effectively fuses the physical constraint of the polarization clue and the data prior in the gradual denoising process, and the surface normal with rich details can still be stably generated in the case of noise interference or information loss. Experimental results show that the method provided by the invention is excellent in surface normal reconstruction in a plurality of complex scenes.
Owner:WUHAN UNIV

Multi-mode-based training method and system for cervical pathology image classification model

The invention relates to the technical field of image classification, in particular to a training method and system of a cervical pathological image classification model based on multiple modes. The method comprises the following steps: acquiring a cervical tissue image and carrying out tissue structure segmentation, forming a nucleus-interstitial-epithelium three-distribution framework, collecting development historical data, confirming a prediction trend of each layer, carrying out environment field simulation through the image, generating a simulated cervical environment field, and carrying out hierarchical evolution prediction on the framework. Evolution mapping images are generated according to the evolution data and classified, finally, a visual basic model is obtained through combined modeling training, image-text fusion is achieved, and a cross-center deployment model system is generated. According to the method, vision-language combined modeling is realized, and the stability and controllability of the whole model structure in image space deformation modeling, semantic cross-modal alignment construction and task-level response flow scheduling are improved.
Owner:GUANGZHOU JINRUI TECHNOLOGY CO LTD

Private weight adaptive heterogeneous data federal cooperative training method and system

The invention provides a private weight self-adaptive heterogeneous data federated cooperative training method and system in the technical field of federated learning and privacy computing, and the method comprises the steps: S1, enabling each client to carry out the differential privacy operation on a local data set based on a private weight, and obtaining a desensitized data set, encoding the desensitized data set through a heterogeneous data encoding model; s2, performing semantic alignment on each coding vector through a contrast learning model to obtain an aligned vector set; s3, training a local model through the alignment vector set, generating a local gradient, extracting local model parameters, and uploading the privacy weight, the local gradient and local difference parameters to a server; and S4, the server trains the global model based on the local difference parameter and the global gradient, extracts the global model parameter and issues the global model parameter to each client for training. The method has the advantages that the compatibility, the flexibility and the efficiency of heterogeneous data federation cooperative training are greatly improved.
Owner:FUJIAN THINKWIN BIG DATA APPLICATION SERVICE CO LTD

Semantic fingerprint adaptive training method for teaching service robot

The invention discloses a semantic fingerprint adaptive training method for a teaching service robot, and relates to the technical field of education neural network real-time training. Comprising six steps of course version semantic fingerprint injection, real-time drift detection and bucket division positioning, small sample correction and high-level weight patching, hierarchical control incremental learning scheduling, shadow reasoning consistency optimization and learning asset registration and cycle verification and tracking. Measuring drifting in real time by using an information entropy self-adaptive window and multi-scale divergence, generating a lightweight weight patch by small sample contrast learning, and performing online loading; shadow channel parallel reasoning is combined with a grading heat exchange superior weight, four-dimensional learning asset tensor is written into a registry through double-clock witness and chain commitment, random sampling verification and singular value performance verification ensure that assets are consistent with robot online examples, and the comprehensive effects of source traceability, risk self-sensing, model self-repairing and compliance full-chain trace reserving are achieved.
Owner:北京爱宾果科技有限公司

STEM teacher intelligent research and repair method and system fusing knowledge graph and graph neural network

The invention relates to the technical field of intelligent education, in particular to an STEM teacher intelligent research and repair method and system fusing a knowledge graph and a graph neural network, and the method comprises an interdisciplinary knowledge graph construction and dynamic updating module which forms a concept association network with timeliness weight through the analysis of multi-source STEM educational resources and the modeling of the graph neural network; the teacher intelligent agent learning companion module is used for converting a teacher request into a teaching scheme with an evidence chain by adopting a thinking chain reasoning mechanism of graph retrieval enhancement and teaching logic constraint; the teacher portrait construction and professional development planning module is used for realizing dynamic quantification of STEM-TPACK (subject teaching knowledge of integration technology) capability characteristics of teachers through multi-modal teaching behavior analysis, and performing joint embedded representation with knowledge graph nodes; and the teacher teaching, learning and research community construction and treatment module constructs an affinity network based on the teacher feature vector, and realizes group intelligent division, self-built large-scale MOOC resource pushing and inter-disciplinary collaborative task generation.
Owner:SHAANXI NORMAL UNIV +1

Large model tool calling hierarchical dynamic optimization method and system based on reinforcement learning

The invention discloses a large model tool calling hierarchical dynamic optimization method and system based on reinforcement learning, and provides a model training mode based on a hierarchical decoupling architecture, which is characterized in that a reward mechanism is adjusted to form a format + tool calling correctness reward, so that the reward efficiency is improved. The correctness rewards are decomposed into three-level verification of names, parameters and values, and formats and correctness reward weights are dynamically adjusted in the training process; thus, the model realizes progressive training from basic structure learning to complex strategy optimization, the generalization ability of the model is enhanced, and fine-grained feedback in the training process is also realized, so that the model can perform gradient updating aiming at specific errors, and the accuracy of the model is improved. The problems of low training efficiency and poor model output accuracy in the traditional technology are avoided; therefore, according to the method, the generalization ability, the training efficiency and the output accuracy of the model are improved, so that the method is very suitable for large-scale application and popularization.
Owner:TIANFU JIANGXI LAB

Artificial intelligence large model training method in heterogeneous multi-machine multi-card environment

The invention discloses an artificial intelligence large model training method in a heterogeneous multi-machine and multi-card environment, and belongs to the technical field of artificial intelligence large model training. Load balancing of heterogeneous equipment is realized by constructing a uniform interface, and the communication efficiency is optimized by adopting hierarchical pipeline aggregation and dynamic quantization compression; the node dynamic adjustment is realized in combination with the elastic topological structure, the problems of poor equipment compatibility, high communication delay and rigid topological structure in the prior art are effectively solved, and the method has the remarkable advantages of improving the utilization rate of heterogeneous computing resources, reducing the cross-node communication overhead and enhancing the fault-tolerant capability of the system.
Owner:SICHUAN HUIXIN INTELLIGENT COMPUTING TECHNOLOGY CO LTD

Federal learning-based privacy protection data sharing and cooperative training method and system

The invention discloses a privacy protection data sharing and cooperative training method and system based on federated learning. The method comprises the steps of receiving software development log data, adaptively judging the sensitivity degree according to a data type, dynamically adjusting noise disturbance intensity according to the sensitivity degree to perform data desensitization, and generating a sensitivity index; selecting a feature extraction strategy, extracting time sequence correlation features from the desensitization data, constructing a dynamic graph structure with a weight, and obtaining a time sequence feature vector through iterative fusion; calculating the time sequence correlation of the time sequence feature vector to obtain a data quality score, and setting a contribution weight based on the quality score to perform parameter aggregation; combining sensitivity indexes with data quality scores to construct a security sharing domain, decoupling global training parameters into knowledge fragments in the domain, formulating a recombination rule, and selectively acquiring the required knowledge fragments by all parties for local training. According to the method, deep collaboration is realized on the premise of protecting data privacy, and the collaboration training effect is improved.
Owner:北京紫荆云科智能技术有限责任公司

Visual encoding method and apparatus, and visual encoding model training method and apparatus

The present application relates to the field of computer vision. Provided are a visual encoding method and apparatus, and a visual encoding model training method and apparatus, which are used for using the same visual encoding model to encode images of different resolutions, and are applied to encoding scenarios for images of more sizes. The visual encoding method comprises: first, acquiring an input image, wherein the input image may be a high-resolution image and may also be a low-resolution image; and then inputting the input image into a visual encoding model, so as to output visual encoding data, wherein the visual encoding model is used for dividing the input image into a plurality of image blocks according to positional embedding, extracting features from each image block, and outputting visual encoding data on the basis of the features of each image block and corresponding positional encoding, the positional embedding is obtained by means of adjusting initial positional embedding on the basis of the difference between the input image and a preset resolution, and the positional embedding may specifically comprise a matrix corresponding to the division of the input image
Owner:HUAWEI TECH CO LTD

Document image tampering detection model training method, tampering detection method and device

The invention provides a training method of a document image tampering detection model and a tampering detection method and device.The training method of the document image tampering detection model comprises the steps that multi-scale visual domain features are extracted from a sample document image, and multi-scale frequency domain compressed sensing features are extracted from frequency domain information; acquiring tampered area edge mask data from the document image; fusing the multi-scale visual domain features and the multi-scale frequency domain compressed sensing features to obtain multi-modal fusion features; performing semi-supervised training on the multi-scale sensing network by taking the multi-scale visual domain feature as a sample feature of a first prediction head, taking the multi-modal fusion feature as a sample feature of a second prediction head, taking a real label or a pseudo label as a sample label and taking joint loss as a loss function to obtain a document image tampering detection model; according to the method provided by the invention, document image tampering pixel-level detection under low labeling cost is realized, and the detection precision of a document image tampering detection model is improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

World model driven decision model training method, system, equipment and product

The invention discloses a world model driven decision model training method, system, device and product, and relates to the technical field of artificial intelligence. According to the scheme, the initial world model is generated through the target video data and the diffusion generation model, and the initial world model is finely adjusted by using three different loss functions, namely the diffusion loss function, the dynamic loss function and the structure maintenance loss function based on the third-order motion prior; physical consistency and high-frequency detail fidelity of short-term and long-range prediction are realized; furthermore, a reward function is automatically generated by using the uncertainty of world model prediction, so that the training efficiency is improved; according to target video data and a world model closed-loop training decision model, collaborative optimization of environment cognition and strategy evolution is realized; and finally, the trained world model and the decision model can be integrated to the target server, closed-loop control of perception-decision-motion execution is realized, the method has low delay, high robustness and expansibility, and the safety of the automatic driving system is improved.
Owner:SHANDONG HAILIANG INFORMATION TECH RES INST

Wafer defect classification method, model training method, system, equipment and medium

The embodiment of the invention provides a wafer defect classification method, a model training method, a system, equipment and a medium. According to the wafer defect classification scheme provided by the invention, the multi-modal test information of the wafer can be acquired, so that a plurality of test maps generated based on the multi-modal test information of the wafer are used as the basis of wafer defect classification, and the test information of different modals (namely, different dimensions) is considered during wafer defect classification; therefore, the accuracy of the classification result can be improved, and an actual manual wafer defect analysis mode can be met. Wherein the plurality of test maps comprise at least two types of maps, and one type of map is generated based on one type of modal test information.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Model training method for health state monitoring of large-scale port equipment

The invention relates to a model training method for health state monitoring of large-scale equipment in a port. The method specifically comprises the following steps: collecting and marking detection data of the large-scale equipment in the port, and dividing a data set; adaptive time-frequency decomposition and feature reconstruction are carried out on the collected data, and an improved adaptive noise complete set empirical mode decomposition and dynamic spectrum kurtosis fusion method is adopted to generate a decomposition and reconstruction matrix; constructing a deep learning model, processing the sample data, and inputting the processed sample data into the model to train the model to obtain a trained model; and collecting detection data of the large-scale equipment in a port in real time, inputting the detection data into the trained model after processing, outputting a health state probability distribution result, and determining a current health state category of the equipment according to a maximum value of probability distribution. According to the method, adaptive time-frequency decomposition and feature reconstruction are carried out on the collected sample data, a deep learning model is constructed and trained, and real-time health detection and degradation evaluation can be carried out on the running state of the equipment.
Owner:YANTAI PORT GRP CO LTD

Large language model training method and device

The embodiment of the invention provides a big language model training method and device, and aims to enable a big language model to have the capability of processing complex services and train the reasoning capability of the big language model. Training is carried out in two stages. In the first stage, a thinking chain is used as supervision fine tuning of a supervision signal, and in the process, the thinking chain can be determined by adopting generation, evaluation and correction modes of a fine-grained single action, so that the flexibility and depth of a reasoning path are improved. The second stage is a reinforcement learning stage, model rewards in the reinforcement learning process comprise correctness rewards and length rewards, the big language model is encouraged to generate a longer and reliable reasoning path, and reward abuse is avoided. According to the scheme, the processing reliability and accuracy of the large language model suitable for complex services can be improved.
Owner:FUDAN UNIVERSITY +1

Robot adaptive training method and device based on reinforcement learning and medium

The invention relates to the technical field of robot training. The robot self-adaptive training method based on reinforcement learning comprises the steps that task sub-target information is generated through a high-level strategy network, the task sub-target information is input into a low-level execution network, an action control instruction is generated according to the task sub-target information, interaction feedback information is collected in the execution process, and the action control instruction is sent to a robot through a robot. Calculating a reward value according to the interaction feedback information, carrying out association processing on the reward value and the scene complexity parameter, executing a dynamic reward shaping operation, generating an adjusted reward signal, generating a strategy model optimized by meta-learning based on the adjusted reward signal, loading the strategy model in a simulation environment, and carrying out dynamic reward shaping. A target strategy model optimized through simulation training is generated, the target strategy model is loaded to the robot, and the robot is controlled to execute task operation in the actual interaction scene. The method has the effect of realizing adaptive task learning of the robot in a multi-interaction scene.
Owner:SEVEN (BEIJING) EDUCATION TECH CO LTD

CT guided liver puncture training method and system based on virtual reality

The invention provides a CT guided liver puncture training method and system based on virtual reality, and relates to the technical field of virtual reality. A deformable liver model and a virtual CT reconstruction engine under respiration driving are constructed, needle body posture mapping and image fusion display are achieved by fusing an inertia-electromagnetic dual-mode sensor, path interaction control, tissue dynamic response and score feedback are supported, the scene difficulty is automatically adjusted based on a training result, and the accuracy and the reliability of the system are improved. Progressive puncture skill training of static breath-holding, shallow breath and free breath scenes is achieved, the sense of reality of training, operation feedback and teaching efficiency are improved, and the system is suitable for development and clinical teaching application of an interventional therapy training system under the guidance of medical images.
Owner:CANCER CENT OF GUANGZHOU MEDICAL UNIV

Model training method, defect detection method and related apparatuses

PCT designated stageWO2025209385A1Image enhancementImage analysisAlgorithmEngineering
Provided in the present application are a model training method, a defect detection method and related apparatuses. The embodiments of the present application can be applied to various scenarios such as computer vision. The model training method trains an initial detection model by means of using a first sampled image set containing some of images that have been used for training and a full newly-added second training image set, and thus, compared with using full historical training data and full newly-added training data to train an initial model, more saves time and reduces the GPU hour consumption; adaptive evaluation of corresponding first weights of model parameters restricts updating of the model parameters with respect to historical training data, so as to solve the problem of knowledge forgetting caused by only using full newly-added data for model fine-tuning, thus improving the learning capability of the detection model; and an optimized detection model obtained by using the model training method is used to detect defects in an image for detection, thus improving the effect and accuracy of defect recognition.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Visual algorithm self-training method based on multi-agent collaborative optimization

The invention discloses a visual algorithm self-training method based on multi-agent collaborative optimization, and the method comprises the following steps: constructing a multi-agent system architecture which comprises a user interaction layer, an intelligent scheduling layer, an A2A protocol communication layer and a professional agent cluster layer; the user interaction layer analyzes a user task intention and generates an execution plan; the scheduling agent calls the professional agent to complete data processing, model construction, training, testing and deployment; a task process is coordinated through a standardized communication mechanism, and task execution is supported by combining an MCP tool set, a knowledge base module and a memory system; and when the task fails, automatically executing rescheduling operation, and finally outputting a self-training result. According to the method, the development efficiency, the self-adaptability and the intelligent level are remarkably improved, and the method is suitable for computer vision tasks such as industrial detection, intelligent security and protection and automatic driving.
Owner:ANHUI HEQING INTELLIGENT ROBOT CO LTD

Large model generation content traceability technology based on model copyright ID watermark embedding

The invention discloses a large model generation content traceability technology based on model copyright ID watermark embedding, which comprises the following steps that: firstly, a large model content generator compiles a dynamic watermark embedding process into an arithmetic circuit, and generates a watermark text and a proof by using a zero-knowledge proof algorithm; then the large model content generator publishes the text with the watermark and the proof together, and hides the copyright ID and the secret key; a large model content user verifies the proof and the watermarked text through a smart contract by adopting a zero-knowledge verification algorithm; after verification is passed, an adversarial sample corresponding to the text with the watermark is generated, an adversarial training method is used for optimizing the text with the watermark, and parameters of the Viterbi balance algorithm are dynamically adjusted and improved. According to the method, the concealment and the generation quality are balanced, high-capacity watermark embedding is realized, the tamper resistance and the accurate traceability of multi-model copyright disputes are improved, and the privacy protection intensity is improved.
Owner:GUANGZHOU UNIVERSITY

Model training method and device

The invention relates to a model training method and device. The method comprises the steps of obtaining a first training image set; performing first fine tuning training on the pre-trained artificial intelligence model by using the first training image set to obtain a preliminary optimization model; acquiring a second training image set; constructing a composite reward function based on the second training image set; wherein the composite reward function is used for performing multi-dimensional quantitative evaluation on the quality output by the model; and performing second fine tuning training on the preliminary optimization model based on the second training image set and the composite reward function to obtain a final optimization model. Therefore, a two-stage differential fine-tuning model strategy is realized, so that the output result of the final optimization model is highly matched with the expectation of a real scene, the problem of insufficient model practicability is fundamentally solved, and the reliability and value of the model after deployment are greatly improved.
Owner:WUHAN KINGSOFT OFFICE SOFTWARE CO LTD +2

Tunnel disease identification model training method and system based on point cloud and image

The invention discloses a tunnel disease recognition model training method and system based on a point cloud and an image, and the method comprises the steps: synchronously collecting tunnel point cloud and image data, carrying out the calibration and registration, achieving the spatial alignment, preprocessing the point cloud, generating a gray-scale image and a depth image, collecting continuous images through a line-scan digital camera, generating a spliced image, and carrying out the recognition of tunnel diseases. Partitioning a large-size image after multi-image space-time synchronization; a multi-branch network is constructed, point cloud geometry and image texture features are extracted by using Point Net / 3DCNN and CNN / Transform respectively, and semantic collaborative fusion is realized through an intermediate layer fusion module; parameters are adjusted according to errors through self-adaptive training, manual labeling dependence is reduced in combination with transfer learning and the like, and convergence is accelerated through online iteration, joint loss and self-adaptive weight to improve generalization; the performance of the model is evaluated through field testing and indexes in various tunnel environments, and the structure is optimized according to data to ensure that engineering is feasible and efficient. The model can accurately identify various diseases on different tasks, and can adapt to complex and changeable working conditions in tunnel detection.
Owner:WUHAN HANNING TECH

Brain tumor imaging diagnosis large model pre-training method, diagnosis method and system

The invention discloses a brain tumor image diagnosis large model pre-training method, diagnosis method and system, and the method comprises the steps: obtaining the image data of a brain tumor patient and a corresponding diagnosis text, and the image data comprises a plurality of sequences; constructing a visual unified model, carrying out complete sequence standard training and missing sequence distillation training on the visual unified model by utilizing the image data, learning unified visual representation of the image data of any sequence combination, and taking characteristics of the complete sequence image data as teacher characteristics in the missing sequence distillation training; the features of the missing sequence image data serve as student features, and distribution alignment of the student features and the teacher features in the feature space is restrained; constructing a visual language model, taking unified visual representation output by the trained visual unified model as input based on a multi-task target, and training the visual language model in combination with the diagnosis text; the trained model can effectively process the sequence missing condition, and the accuracy and robustness of diagnosis are improved.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1

Intelligent digital human training method and system based on multi-modal interaction

The invention discloses an intelligent digital human training method and system based on multi-modal interaction, and belongs to the technical field of semantic indexing.The method specifically comprises the steps that voice, vision and text data are analyzed and converted into high-dimensional feature vectors through a modal exclusive encoder, the high-dimensional feature vectors are projected to a unified semantic space through a cross-modal semantic mapping model, and the high-dimensional feature vectors are obtained; generating a semantic primitive containing a modal identifier, a core semantic tag and a feature weight; semantic primitives are used as nodes, directed edges and edge weight table association strength are established based on semantic similarity, typical scene node connection weights are strengthened, and a mesh map containing intra-modal hierarchy and inter-modal cross association is formed; constructing a double-layer index on the basis of the mesh map; semantic primitives are extracted from newly added data, the position of a new node in an association graph is determined through a graph matching algorithm, an association edge with an existing node is automatically established, and a lower-layer modal exclusive index is synchronously updated.
Owner:JIANGXI INST OF FASHION TECH

Three-dimensional target detection model training method and device based on image-guided depth completion and multi-stage iterative fusion

The invention discloses a multi-modal three-dimensional target detection method and device based on image-guided depth completion and multi-stage iteration fusion, and the method comprises the steps: firstly, predicting a dense depth map through an image-guided depth completion module by using the context information of an image, and carrying out the image-guided depth completion; the depth map is fused with a sparse depth map generated by the laser radar in a mask guiding manner, so that a high-quality complemented depth map is generated, and the accuracy of subsequent view angle conversion is improved; and then, through a multi-stage iterative fusion module, iterative fine-grained fusion is carried out on the converted image aerial view features and point cloud aerial view features, so that modal conflicts are effectively relieved, and the expression ability of fusion features is enhanced. According to the invention, through accurate depth information completion and efficient multi-modal feature fusion, the precision and robustness of three-dimensional target detection can be significantly improved, and especially the effect is more obvious when a long-distance target or a blocked target and other difficult targets are processed.
Owner:ZHEJIANG COLLEGE OF ZHEJIANG UNIV OF TECHOLOGY

Target detection method and device, model training method and device, electronic equipment and medium

The invention relates to the technical field of data processing, and provides a target detection method and device, a model training method and device, electronic equipment and a medium. The target detection method comprises the steps that a to-be-recognized image and a query text are acquired, and the query text is used for querying a target object corresponding to the query text in the to-be-recognized image; performing image recognition on the to-be-recognized image to obtain image description features and region detection visual features; performing regional multi-modal fusion processing on the image description features and the regional detection visual features to obtain regional multi-modal fusion features; performing feature fusion processing on text features obtained based on the query text and the regional multi-modal fusion features to obtain text regional fusion features corresponding to the query text; and a target detection result is obtained based on the text features and the text region fusion features, so that the fusion degree of text semantics and image region features is improved, and the accuracy and robustness of target detection in a complex scene are improved.
Owner:BEIJING JIZHI DIGITAL TECH CO LTD

Medical report generation method, model training method, equipment and medium

The invention discloses a medical report generation method, a model training method, equipment and a medium, and the model training method comprises the steps: constructing a medical report generation model framework which comprises a global semantic collaborative multi-modal enhancement module, a visual encoder, a text encoder, a medical insight analyzer and an LLM decoder; wherein the global semantic collaborative multi-modal enhancement module respectively enhances a medical image and a medical report by utilizing a selected image enhancement strategy and a text enhancement strategy, and the medical insight analyzer comprises a fine-grained structure learning device and a global context guide learning device which are connected in sequence so as to enhance the cross-modal alignment capability; and performing intelligent collaborative optimization by taking a strategy set formed by an image enhancement strategy and a text enhancement strategy and architecture configuration parameters of the medical insight analyzer as optimization targets to obtain an optimal medical report generation model. The medical report generation performance can be effectively improved.
Owner:CENT SOUTH UNIV

Database question and answer model training method and device, storage medium and computer equipment

The invention discloses a database question and answer model training method and device, a storage medium and computer equipment, and the method comprises the steps: associating standard structured query language statements, standard execution result answers and standard natural language questions, and generating training annotation data; and collecting simulation derivation problems possibly proposed for the database to obtain non-labeled data for training. Based on a GRPO reinforcement learning framework and a scoring reward function provided by a double-tower model, training is carried out on the scoring reward function by utilizing training labeling data, supervised fine tuning training is carried out on a database question and answer model, and non-labeling data for training, format rewards, executable rewards and scoring rewards of the scoring reward function are combined, so that the scoring reward function of the database question and answer model is obtained. And continuing to train the database question and answer model after supervised fine tuning training. Preliminary training is carried out through a small amount of annotation data, then subsequent training is carried out through non-annotation data, the reasoning ability of the model can be stimulated, the annotation cost is reduced, and the training efficiency is improved.
Owner:SHENZHEN QIANHAI HUANRONG LIANYI INFORMATION TECHNOLOGY SERVICES CO LTD

Streaming knowledge injection and adversarial self-optimization large language model training method and system

The invention discloses a streaming knowledge injection and adversarial self-optimization large language model training method and system, and the method comprises the steps: collecting the newest knowledge of an authoritative information source in real time, converting the newest knowledge into a structured constraint rule through a semantic analyzer, and dynamically updating a knowledge base; based on the updated knowledge base, adopting a PPO algorithm to optimize a generator, actively constructing a high-risk adversarial sample, and forcing the model to expose security vulnerabilities; after user input and adversarial samples are input into the model, total loss is calculated through constraint detection, gradient updating is blocked if the total loss exceeds a threshold value, and otherwise, a multi-level safety verification stage is started; and finally, fusing a security verification result and the adversarial loss, updating a total loss function and cooperatively adjusting model parameters to form a continuous self-optimization training cycle. According to the method, compliance is guaranteed through streaming knowledge injection, vulnerabilities are actively mined in combination with adversarial training, verification precision is improved by means of multi-level detection and dynamic threshold adjustment, and safety and reliability of a large language model in a complex scene are enhanced.
Owner:LANZHOU UNIV