Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

63 results about "Preference learning" patented technology

Preference learning is a subfield in machine learning, which is a classification method based on observed preference information . In the view of supervised learning, preference learning trains on a set of items which have preferences toward labels or other items and predicts the preferences for all items.

Indoor three-dimensional image intelligent rendering method based on image processing

The invention relates to the technical field of image processing, and provides an indoor three-dimensional image intelligent rendering method based on image processing. Comprising the following steps of indoor multi-view image acquisition, image preprocessing, intelligent analysis of indoor scene elements, construction of an indoor three-dimensional initial model with attributes, intelligent generation of rendering parameters, adaptive LOD real-time rendering and intelligent interaction optimization. According to the intelligent interaction optimization system, through deep fusion of natural language processing, user preference learning and real-time rendering technologies, a set of efficient, visual and personalized virtual scene rendering interaction process is constructed. According to the method, the problems that traditional graphic software is complex in operation, high in learning cost and low in debugging efficiency are solved, and the user satisfaction and creation efficiency are remarkably improved through an intelligent recommendation and rapid iteration mechanism.
Owner:贵州轻工职业大学 +1

Systems and methods for reinforcement learning networks with iterative preference learning

Embodiments described herein provide a reinforcement learning framework for neural network models to generate outputs that align with desired human preference. In at least one embodiment, cross-prompts are generated from an original prompt to elicit a response from the neural network model.
Owner:SALESFORCE INC

Layering and preference learning-based personalized dynamic treatment strategy generation system and method for Parkinson's disease

The invention discloses a layering and preference learning-based personalized dynamic treatment strategy generation system for Parkinson's disease, which comprises a dynamic state and preference representation module, a preference-based reward function learning module, a layering treatment strategy generation module and a treatment path simulation and interactive output module, the system learns preferences from doctor-patient interaction and generates layered, dynamic and explainable personalized long-term treatment strategies, so that making of personalized long-term treatment decisions is achieved, and the problems that in an existing Parkinson's disease reinforcement learning scheme, an award function is fixed, and long-term planning ability is lacked are solved.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Intelligent sound field adaptive system of digital professional sound equipment

The invention belongs to the field of artificial intelligence, particularly relates to an intelligent sound field adaptive system of digital professional sound equipment, and aims to solve the problems of inaccurate sound field regulation and control, response lag, dependence on manual tuning and the like in a complex acoustic environment. The system comprises a sound field sensing module, an acoustic modeling and analysis module, a self-adaptive sound field regulation and control engine, a multi-channel digital audio processing unit and a feedback optimization module, and dynamic sound field modeling, real-time audio processing and environment self-adaptive regulation and control are realized through distributed sensing, hybrid modeling, multi-target optimization and closed-loop feedback. The system supports rapid re-calibration, multi-scene memory and user preference learning, ensures voice clarity and music fidelity, balances full-field hearing consistency, and significantly reduces manual intervention requirements.
Owner:深圳市多乐声电子有限公司

Automatic control system and control method for sunshade

The invention relates to the technical field of intelligent sunshade equipment control, in particular to an automatic control system and method for a sunshade, and aims to solve the problem that an existing automatic control system for the sunshade generally does not fully integrate multi-dimensional environment parameters and user intervention behaviors in the aspect of strategy generation, and the strategy generation efficiency is high. The problems that energy consumption, comfort and illumination uniformity are difficult to optimize synchronously under a dynamic meteorological condition, so that a control strategy is lack of individuation and foresight are solved; multi-dimensional environment semantic modeling, short-term illumination trend prediction and user preference learning are fused through a sun shading strategy dynamic generation module, a dynamic weighted multi-target optimization mechanism is constructed, energy consumption, comfort and illumination uniformity are synchronously considered under the condition that equipment constraints are met, a control target and an evaluation standard are updated in real time based on user intervention records, and the system is high in practicability. Accurate, prospective and highly personalized sunshade strategy generation is realized, and the intelligent level and comprehensive performance of the system are remarkably improved.
Owner:ZHEJIANG SAIOU SUNSHADE TECH CO LTD

Key-value memory network-based active recommendation method for design knowledge of complex mechatronic systems

ActiveCN115859822BImprove and refine performanceWeaken barriers to information exchangeDesign optimisation/simulationNeural architecturesSystem design processNetwork output
The application discloses a kind of based on key value memory network's complex electromechanical system design knowledge active recommendation method.First, the scene feature semantic information extraction of software platform log file in the design process of complex electromechanical system is carried out, and scene ontology is established based on scene feature semantic information, and system scene knowledge base is formed by scene ontology and original knowledge base;Then, the scene maximum frequent sequence is used to describe the scene sequence feature similarity between designers;Again, the knowledge item interaction sequence of all designers is learned, and the knowledge item sequence preference vector corresponding to all designers is obtained;Further, input to key value memory network, and the initial design knowledge active recommendation sequence is obtained;Finally, after knowledge item selection, the final design knowledge active recommendation sequence of each designer is obtained.The application obtains design knowledge active recommendation sequence more in line with the needs of designers, and then improves the efficiency of complex electromechanical system design.
Owner:HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS

Large language model lifelong alignment method based on memory enhancement

The invention discloses a memory enhancement-based large language model lifelong alignment method. The method comprises the steps of focus preference optimization and short-time to long-time memory consolidation. The focus preference optimization adaptively focuses the learning focus on a new preference sample or a preference sample with an uncertain model through an improved preference learning loss function, meanwhile, the updating amplitude of fully learned knowledge is reduced, and a historical alignment result is protected while new preference is learned; the short-term to long-term memory consolidation is used for simulating a human memory mechanism, denoising is performed on short-term parameter update through singular value decomposition, an update part conflicting with past knowledge is identified and suppressed by projecting to a historical knowledge subspace, and finally refined conflict-free knowledge is integrated into long-term parameters of a model. According to the method, knowledge of the model is effectively accumulated and reserved in continuous and diversified alignment tasks, catastrophic forgetting is remarkably inhibited, and the stability, reliability and alignment consistency of the model in a dynamic and long-term deployment environment are improved.
Owner:EAST CHINA NORMAL UNIV +1

Security basic large model based on preference reinforcement learning and training method

The invention relates to the technical field of model training, and particularly discloses a security basic large model based on preference reinforcement learning and a training method, and the model comprises a processing module which is used for obtaining all positive samples and all negative samples of an input video group based on the portrait description of all videos of the input video group; the training module is used for obtaining all comparison sample pairs based on all positive samples and all negative samples of the input video group and a preset unlabeled video group; the calculation module is used for obtaining an iterative sequence of all the comparison sample pairs based on all the comparison sample pairs and a preset preference learning technology; and the iteration module is used for obtaining an iteration screening straight line of the security and protection basic large model based on the iteration sorting of all the comparison sample pairs so as to obtain the security and protection basic large model. According to the invention, a model which can be trained by using a small amount of training data and can accurately derive correct human portrait description of an input video is realized.
Owner:BEIJING QINGSI INTELLIGENT TECHNOLOGY CO LTD

A niche preference learning method for generating content based on user text

The application discloses a kind of based on user text generation content's minority preference learning method in the field of information retrieval, comprising the following steps: data preprocessing operation is carried out to the user text generation content obtained;The data obtained by pre-processing establishes a hierarchical Bayesian model, obtains joint distribution model;Model parameters are learned by Gibbs sampling method, and the formula of mass preference distribution and minority preference distribution is obtained;The meaning of the user minority preference based on user text generation content is analyzed using the learned model parameters;The target user under the minority preference is found using the user minority preference distribution.The method of the application distinguishes mass preference and minority preference from the perspective of user preference, identifies the specific meaning of user minority preference using the good interpretability of hierarchical Bayesian method, provides the opportunity for small and medium-sized enterprises to enter suitable niche market, and the minority preference distribution of each user is beneficial to the enterprise to find out the target user of relevant niche market.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Linear dimming type mixed light illumination method and system based on AI fuzzy control

The invention discloses a linear dimming type mixed light illumination system and method based on AI fuzzy control, and belongs to the technical field of LED illumination, the system collects information such as ambient brightness, time, human body activity and the like in real time through an environment sensing module, processes data through an AI-driven fuzzy control algorithm, and outputs target color temperature and brightness level through reasoning. And then, accurate mixed light output is realized through color temperature and brightness coordinate conversion and linear dimming driving. In addition, the system also adopts a chromaticity stability feedback and user preference learning mechanism to continuously optimize the dimming and toning effects. The problems of safety problem, light color drift and lack of intelligent adjustment capability caused by stroboflash of digital dimming in the prior art are solved. By introducing a linear dimming mode and applying an AI fuzzy control algorithm, the system can perform dynamic adjustment according to environmental changes and automatically select the most suitable dimming and toning strategy, thereby ensuring that accurate and stable color temperature and brightness are output in different environments.
Owner:江西省通讯终端产业技术研究院有限公司

An online learning behavior personalized recommendation system based on cluster analysis

The present application patent belongs to the field of big data analysis technology and personalized recommendation system, and particularly relates to an online learning behavior personalized recommendation system based on clustering analysis method, which mainly comprises four functional modules, namely an online learning behavior data information management module, a data analysis modeling module, a student learning portrait display module and a same-type learning friend personalized recommendation module. The system firstly collects and cleans data from an online learning platform, constructs an online learning student portrait label system, and based on the clustering analysis method, creates an online student portrait for each online student from three dimensions of learning attitude style, learning interest preference and learning level ability. Meanwhile, based on the same-type learning friend personalized recommendation model, students with similar online learning behaviors are recommended to each other to further stimulate the learning enthusiasm of students through mutual exchange and discussion.
Owner:CHINA UNICOM (SHANGHAI) IND INTERNET CO LTD

Low-impact development facility optimization layout system based on multi-objective optimization and uncertainty analysis

PendingCN121328289AEnsemble learningBiological modelsData acquisitionLow-impact development
The invention belongs to the technical field of urban rainwater management, and particularly relates to a low-impact development facility optimization layout system based on multi-objective optimization and uncertainty analysis. The system comprises a data acquisition and preprocessing layer, an uncertainty modeling and analysis layer, a multi-target collaborative optimization layer, an intelligent decision support layer and a dynamic adaptation and optimization layer. The data acquisition and preprocessing layer serves as a basic support layer of the system and is responsible for acquisition, cleaning, standardization and integration of multi-source heterogeneous data; the uncertainty modeling and analysis layer serves as a core theory layer of the system and establishes a ternary integrated uncertainty processing framework fusing fuzzy logic, the probability theory and machine learning; the multi-objective collaborative optimization layer serves as an algorithm engine layer of the system and integrates a hybrid optimization technology of an adaptive evolutionary algorithm, machine learning and heuristic search; the intelligent decision support layer serves as an interactive interface layer of the system and provides personalized decision support service based on user preference learning; and the dynamic adaptation and optimization layer is used as a self-learning layer of the system and realizes online optimization and adaptive adjustment based on real-time monitoring data. The invention provides a low-impact development facility optimization layout system based on multi-objective optimization and uncertainty analysis.
Owner:CHINA MCC5 GROUP CORP LTD

Code repairing system and method based on preference learning

The invention discloses a preference learning-based code repair system and method. A code repair enhancement module of the system trains the large language model based on the code data set so as to construct an initial code repairer; a repair preference data generation module generates candidate repair codes from the programming tasks and the error codes through an initial code repairer, and constructs a preference learning data set after preference data selection; the code preference learning module is used for training the initial code restorer through the preference learning data set so as to obtain a final code restorer; and the code repair generation module generates a repair code with a small modification range for the to-be-repaired code data through the final code repairer. According to the method and the device, the repair code which is accurately repaired and has a smaller modification range can be intelligently generated according to actual requirements in a development scene, so that code problems are quickly positioned and solved, unnecessary modification of the code is effectively reduced, the consistency and readability of the code are kept, the time cost of manual debugging is reduced, and the development efficiency is improved.
Owner:COMPUTER INNOVATION TECH RES INST OF ZHEJIANG UNIV +2

Hardware autonomous control method and system based on spatial intelligence and self-evolution learning

PendingCN122284358AEvolutionary learningLinguistic model
This invention discloses a hardware autonomous control method and system based on spatial intelligence and self-evolutionary learning, an electronic device, and a computer-readable storage medium. The system includes: a multi-source device discovery module for automatically scanning intelligent hardware devices and establishing a unified device model through multiple communication protocols; a capability reflection module for automatically extracting device control capabilities and parameter constraints from protocol metadata; a skill management module for storing and loading skill packages and establishing a mapping index from device type to skill package; a rule engine module for performing millisecond-level deterministic evaluation of sensor data and generating control commands; an intelligent decision engine module for orchestrating planning nodes and execution nodes based on state diagrams and making context-aware control decisions through a large language model; a command security module; an execution verification module; a multi-layer memory module; a preference learning module; and a multi-layer self-evolutionary engine module.
Owner:FULAI DIGITAL (BEIJING) INTELLIGENT TECHNOLOGY CO LTD

Construction method of mixed reward model for automated radiotherapy plan

The invention relates to a method for constructing a mixed reward model for an automatic radiotherapy plan, and the method comprises the steps: defining the overall architecture of the mixed reward model, which comprises a clinical target quantification module, an ideal DVH evaluation module and a human expert preference learning module; constructing a clinical target quantification module, obtaining the current treatment plan, and performing quantitative evaluation on the achievement condition of the current treatment plan on each clinical index in combination with a preset clinical treatment scheme; constructing an ideal DVH evaluation module, predicting an ideal DVH parameter value possibly reached by the current patient, and comparing the ideal DVH parameter value with an actual DVH parameter of the current plan; constructing a human expert preference learning module, collecting preference data of experts on the treatment plan, extracting visual features from the dose distribution image, and training a preference model; and determining a generation method of the comprehensive reward signal, and generating and outputting the comprehensive reward signal. The method provides a high-robustness evaluation mechanism for radiotherapy plan automatic optimization.
Owner:SHANGHAI BUSINESS SCHOOL

Context preference learning method, device and equipment based on large language model

The invention relates to a context preference learning method, device and equipment based on a large language model. The method comprises the following steps: configuring a task environment and defining performance indexes; automatically generating multiple groups of initial reward functions through a large language model; training a plurality of reinforcement learning agents in parallel and collecting behavior data; automatically identifying an optimal reward function and a worst reward function based on a weighted scoring mechanism; and driving a large language model to generate an improved function in combination with comparison difference information, and performing iterative optimization. According to the scheme, the bottleneck of manual design can be broken through, and dynamic weight adjustment and cross-scene generalization of the reward function are realized; the decision-making efficiency and safety are obviously improved in the scenes of customer service, industrial control and the like; the labor cost is reduced through unsupervised preference evaluation, and continuous adaptive optimization of the strategy model is supported.
Owner:深圳市和讯华谷信息技术有限公司

An occupant personalized preference driven intelligent chassis control method, system, device and storage medium based on a vision-language-action model

This invention discloses a method, system, device, and storage medium for intelligent chassis control driven by occupant personalized preferences based on a vision-language-action model, belonging to the field of intelligent vehicle control and intelligent chassis collaborative control technology. The method includes: collecting multi-source information; constructing a personalized dynamic experience preference profile of the occupant; performing scene-preference multimodal joint modeling to form a unified scene-preference semantic representation; inputting the vision-language-action model to generate personalized chassis style intent; generating dynamic performance constraints and controller parameter adjustment targets; generating specific chassis control targets; executing multi-actuator personalized collaborative control; and performing parameter recovery and preference learning updates. This method realizes vision-language-action joint reasoning between occupant personalized preferences, visual scene semantics, and chassis control actions, improving the overall performance of autonomous and assisted driving vehicles in terms of comfort, stability, safety, and personalized experience.
Owner:RUIXING INTELLIGENT (YANCHENG) TECH CO LTD

System

The system according to the embodiment aims to simplify the process of travel planning and to make personalized suggestions according to the user's preferences.SOLUTION: A system according to an embodiment includes a user preference learning unit, a personalization proposal unit, a real-time information acquisition unit, an automatic budget control unit, a collective reservation unit, and a customization unit. The user preference learning unit learns the user's preference. The personalized suggestion unit suggests tourist spots, restaurants, and activities personalized based on the user's preference. The real-time information acquisition unit acquires real-time information of a travel destination. The automated budget manager optimizes the travel plan based on the user's budget. The collective reservation unit makes a collective reservation based on the travel plan. The customization unit customizes the travel plan in accordance with a user's request.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Building facade control method and system based on network fine tuning, controller and storage medium

The invention relates to the technical field of building control, and discloses a building facade control method and system based on network fine tuning, a controller and a storage medium, and the method comprises the steps: carrying out the standardization processing of real-time environment data, and obtaining an environment state vector; inputting the environment state vector into an initial controller, controlling an external facade component and receiving feedback data of a target user; constructing a preference pair in the current period according to the feedback data so as to optimize the initial controller to obtain a target controller; and inputting the environment state vector of the target area in the next period into a target controller so as to adjust the external facade component. According to the method, through deep reinforcement learning driven by deep fusion data and a lightweight online preference learning technology, efficient management of external facade components is improved, and the adaptive capacity, the individuation degree and the multi-target real-time optimization efficiency of equipment are improved.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Tool integration reasoning method based on self-evolution preference learning

The invention provides a tool integrated reasoning method through self-evolution preference learning. Comprising two parts: 1, training data construction based on information entropy; 2, a multi-stage self-evolution training normal form; the training data construction process based on the information entropy is used for constructing training data and comprises a data source screening process and an entropy guide sampling process; the multi-stage self-evolution training normal form is used for improving the tool integration reasoning capability of the model, and comprises a two-stage normal form of supervision fine tuning and self-evolution direct preference optimization; according to the reasoning method, for a given input instruction I and a large language model theta, a final answer y is generated through a calculation process. By means of the technical scheme, the effects that efficiency and accuracy are improved, the reasoning process is more stable and simpler, the training data construction cost is lower, and the model generalization ability is higher can be achieved.
Owner:RENMIN UNIVERSITY OF CHINA

An AI customer service interaction strategy self-adaptive adjustment method based on reinforcement learning

PendingCN122332520AUncertainty representationSelf adaptive
This invention discloses an AI customer service interaction strategy adaptive adjustment method based on reinforcement learning, comprising the following steps: acquiring multi-source interaction data and constructing a state-action trajectory input structure; combining human feedback on customer service interaction quality preferences to construct trajectory preference comparison samples and preference condition vectors; constructing a human preference learning framework; constructing a conditional preference reward generation structure within the human preference learning framework to generate trajectory preference scores and preference uncertainty representations; constructing a preference attribution credit allocation structure to generate dense reward sequences corresponding to interaction rounds; performing adaptive preference updates based on online interaction data; constructing a reinforcement learning strategy optimization process to generate an adaptive strategy model; and using the adaptive strategy model to execute online interactions and iteratively update the learning process. This invention achieves fine-grained optimization, stable updates, and continuous adaptive adjustment of customer service strategies.
Owner:TUYOU NETWORK TECHNOLOGY (SHENYANG) CO LTD

Human-machine cooperation method for solving human deviation

PendingCN121766465Ainhibitory dominance effectReduce the number of preference queriesMathematical modelsMachine learningAlgorithmPreference relation
The invention discloses a man-machine cooperation method for solving human deviation. The man-machine cooperation method comprises the steps that initialization is carried out; iteratively executing batch Thompson sampling, batch Thompson sampling, preference query and data updating and Gaussian process posteriori updating until the maximum number of iterations is reached; and after iteration is finished, returning an optimal action corresponding to the maximum potential function value in the action space. According to the embodiment, the long-tail preference relationship problem is fundamentally solved. In the batch Thompson sampling stage, diversified candidate preference pairs are generated through an adaptive covariance scaling factor and a double-independent sampling mechanism; in the sub-mode marginal gain evaluation stage, the dominant effect of head preference is effectively inhibited by utilizing the profit decreasing characteristic and marginal gain maximization of a sub-mode function; and in a Gaussian process posteriori updating stage, the observed preferences are integrated into a Bayesian framework, and posteriori distribution is refined to guide subsequent sampling. By optimizing the preference learning process, the preference query times are remarkably reduced, and the learning efficiency and accuracy are improved.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

System

An object of a system according to an embodiment is to provide information personalized on the basis of an action history or hobbies and preferences of a user in real time.SOLUTION: A system includes a collection unit, a learning unit, a generation unit, and a provision unit. The collection unit collects an action history or hobbies and preferences of a user. The learning unit learns the data collected by the collection unit. The generation unit generates personalized information on the basis of the data learned by the learning unit. The provision unit provides the information generated by the generation unit to the user in real time.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Wood texture aesthetic feature extraction and personalized matching system

InactiveCN121505583ACharacter and pattern recognitionCommerceFeature extractionComputational aesthetics
The invention discloses a wood texture aesthetic feature extraction and personalized matching system, which belongs to the technical field of wood texture analysis and calculation aesthetics and comprises an image acquisition module, an aesthetic feature extraction module, a preference learning module, a similarity calculation module, a personalized recommendation module and a texture continuity optimization module. The method comprises the following steps: quantitatively extracting multi-dimensional aesthetic characteristic parameters of wood through multi-scale characteristic analysis and visual perception, constructing a personalized preference model based on a user selection history, calculating weighted similarity between to-be-matched wood and user preference, generating a personalized recommendation list, and providing an optimal arrangement scheme for splicing multiple pieces of wood. The technical problems that in the prior art, wood texture aesthetic feature quantification is not comprehensive, a personalized recommendation mechanism is lacked, and the multi-piece splicing effect is poor are solved, and the method is suitable for the fields of high-end furniture manufacturing, interior design, wood product personalized customization and the like.
Owner:FUJIAN FORESTRY VOCATIONAL TECH COLLEGE

Network strategy optimization method and system based on user behavior and preference learning

The invention relates to the field of network communication, and provides a network strategy optimization method and system based on user behaviors and preference learning. The method comprises the following steps: collecting strategy operation behaviors and network quality sensing data of a user to obtain a user behavior data set; inputting the user behavior data set into a machine learning model for training so as to learn an association relationship between user behavior characteristics and network quality indexes, and obtaining a user preference weight vector; obtaining a quality parameter and available network strategies of the current network environment, and performing quantitative evaluation on the plurality of available network strategies according to the user preference weight vector to obtain a comprehensive score of the plurality of available network strategies; and performing priority ranking on the available network strategies based on the comprehensive score, and selecting a target network strategy according to a strategy availability state and a score result. According to the method, the personalized preference of the user on the network quality index can be learned, intelligent evaluation and dynamic switching of the network strategy are realized, and the user experience is improved.
Owner:E SURFING VISION TECHNOLOGY CO LTD

Federal multi-modal preference learning and alignment method for mode-lacking condition

The invention discloses a federal multi-modal preference learning and alignment method for a mode lacking condition, and aims to solve the problem of unstable alignment caused by image-text data mode lacking, preference noise and cross-domain conflict in multi-client privacy protection. According to the method, a modal state is identified and a conditional preference pair is constructed at a client: cross-modal consistency scoring is carried out to screen good and bad responses when images and texts are complete, a visual attribute lexicon is introduced to punishment suppression assume when the texts are lack of texts, and entity existence verification punishment illusion is carried out when the images are lack of texts; noise updating is reduced in combination with sample-level and client-level two-level gating based on prediction uncertainty; and the server side forms multi-preference experts by utilizing low-dimensional preference signature clustering and issues and fuses the multi-preference experts according to similarity routing. According to the method, the consistency and robustness of multi-modal generation are improved in a modal-lacking scene, and negative migration and illusion are reduced.
Owner:GUANGXI NORMAL UNIV

Interactive multi-objective Bayesian optimization method and system based on dynamic preference learning

The invention relates to the technical field of artificial intelligence optimization, and relates to an interactive multi-objective Bayesian optimization method and system based on dynamic preference learning. The method comprises the following steps: constructing a Gaussian process regression model for each objective function, and training by using an initial data set; performing weighted Euclidean distance non-dominated sorting on the non-dominated solutions by using an NSGA-II algorithm to obtain a candidate solution set; selecting a solution pair with the highest information gain from the candidate solution set, and initiating pairwise preference inquiry to a decision maker to obtain preference feedback; sampling by using a Markov chain Monte Carlo method and updating a preference weight; constructing a weighted Chebyshev utility function, and triggering a self-switching acquisition strategy when invalid utility in a window is promoted; and carrying out real objective function evaluation and iterative updating. According to the method, the preference information of the decision maker can be dynamically embedded in the multi-objective optimization process, and the optimal solution most conforming to the preference can be efficiently obtained with fewer function evaluation times.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

An interpretable reward model construction method based on a sparse self-encoder

The application provides an interpretable reward model construction method based on a sparse self-encoder, and the method comprises the following steps: constructing and pre-training a sequence-level sparse self-encoder, and constructing an interpretable reward model, wherein the encoder part of the pre-trained sparse self-encoder is integrated after the first layer of a basic language model, and all original layers after the first layer of the basic language model are discarded; the reward model is trained based on preference data, wherein the encoder parameters of the sparse self-encoder are frozen, the parameters of the front layers of the basic language model and the weights of a linear value head are trained, the weights of the linear value head are optimized by using a pair of preference data through a preference learning loss function, and the model learns the contribution weight of each interpretable feature to human preference. The method can be simply and efficiently applied to various real scenes, and is particularly suitable for reinforcement learning training of a strategy model under dynamic preference and large-scale data filtering.
Owner:UNIV OF SCI & TECH OF CHINA

An ultrasonic tongue silent speech recognition method based on modal transfer learning

The application discloses an ultrasonic tongue silent speech recognition method based on modal transfer learning, relates to the technical field of speech information processing, and proposes a multi-stage and multi-modal modeling pipeline for two main challenges of SSR, namely, a limited data set and a confusing silent signal. Four-step processes are introduced to avoid overfitting and two additional modules are used to extract multi-modal information. Based on a joint training strategy, modal preference learning (MPL) is proposed to promote the optimization of the target task by utilizing cross-modal prior knowledge.
Owner:THE ACAD OF TIANJIN UNIV HEFEI +1

Cross-domain sequence recommendation method based on causal inference and preference evolution

The invention belongs to the technical field of cross-domain sequence recommendation, and discloses a cross-domain sequence recommendation method based on causal inference and preference evolution. Through a cue word template designed based on cross-domain co-occurrence frequency, article semantic information is enhanced by using a large language model. Performing intra-domain sequence preference learning through the domain-specific interest evolution model to obtain intra-domain preferences; and performing cross-domain sequence preference learning through a time sequence-domain dual-condition hybrid expert mechanism to obtain cross-domain preference. Causal depolarization is designed to decouple the user intra-domain preference and the user cross-domain preference from the user intra-domain activeness and the user cross-domain activeness, recommendation is generated, and the accuracy of a cross-domain recommendation result is improved. According to the method, by means of a causal enhanced preference learning mechanism, the limitation of an existing method is effectively solved, and the most advanced performance is achieved in single-domain and cross-domain sequence recommendation benchmark tests.
Owner:NORTHEASTERN UNIV CHINA