Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
46 results about "Preference learning" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
Preference learning is a subfield in machine learning, which is a classification method based on observed preference information . In the view of supervised learning, preference learning trains on a set of items which have preferences toward labels or other items and predicts the preferences for all items.
The invention relates to the technical field of image processing, and provides an indoor three-dimensional image intelligent rendering method based on image processing. Comprising the following steps of indoor multi-view image acquisition, image preprocessing, intelligent analysis of indoor scene elements, construction of an indoor three-dimensional initial model with attributes, intelligent generation of rendering parameters, adaptive LOD real-time rendering and intelligent interaction optimization. According to the intelligent interaction optimization system, through deep fusion of natural languageprocessing, user preference learning and real-time rendering technologies, a set of efficient, visual and personalized virtual scene rendering interaction process is constructed. According to the method, the problems that traditional graphic software is complex in operation, high in learning cost and low in debugging efficiency are solved, and the user satisfaction and creation efficiency are remarkably improved through an intelligent recommendation and rapid iteration mechanism.
The invention discloses a layering and preference learning-based personalized dynamic treatment strategy generation system for Parkinson's disease, which comprises a dynamic state and preference representation module, a preference-based reward function learning module, a layering treatment strategy generation module and a treatment path simulation and interactive output module, the system learns preferences from doctor-patient interaction and generates layered, dynamic and explainable personalized long-term treatment strategies, so that making of personalized long-term treatment decisions is achieved, and the problems that in an existing Parkinson's diseasereinforcement learning scheme, an award function is fixed, and long-term planning ability is lacked are solved.
The invention belongs to the field of artificial intelligence, particularly relates to an intelligent sound fieldadaptive system of digital professional sound equipment, and aims to solve the problems of inaccurate sound field regulation and control, response lag, dependence on manual tuning and the like in a complex acoustic environment. The system comprises a sound field sensing module, an acoustic modeling and analysis module, a self-adaptive sound field regulation and control engine, a multi-channel digital audioprocessing unit and a feedback optimization module, and dynamic sound field modeling, real-time audio processing and environment self-adaptive regulation and control are realized through distributed sensing, hybrid modeling, multi-target optimization and closed-loop feedback. The system supports rapid re-calibration, multi-scene memory and user preference learning, ensures voice clarity and music fidelity, balances full-field hearing consistency, and significantly reduces manual intervention requirements.
The invention relates to the technical field of intelligent sunshade equipment control, in particular to an automatic controlsystem and method for a sunshade, and aims to solve the problem that an existing automatic controlsystem for the sunshade generally does not fully integrate multi-dimensional environment parameters and user intervention behaviors in the aspect of strategy generation, and the strategy generation efficiency is high. The problems that energy consumption, comfort and illumination uniformity are difficult to optimize synchronously under a dynamic meteorological condition, so that a control strategy is lack of individuation and foresight are solved; multi-dimensional environment semantic modeling, short-term illumination trend prediction and user preference learning are fused through a sun shading strategy dynamic generation module, a dynamic weighted multi-target optimization mechanism is constructed, energy consumption, comfort and illumination uniformity are synchronously considered under the condition that equipment constraints are met, a control target and an evaluation standard are updated in real time based on user intervention records, and the system is high in practicability. Accurate, prospective and highly personalized sunshade strategy generation is realized, and the intelligent level and comprehensive performance of the system are remarkably improved.
The invention discloses a linear dimming type mixed light illumination system and method based on AI fuzzy control, and belongs to the technical field of LED illumination, the system collects information such as ambient brightness, time, human body activity and the like in real time through an environment sensing module, processes data through an AI-driven fuzzy control algorithm, and outputs target color temperature and brightness level through reasoning. And then, accurate mixed light output is realized through color temperature and brightness coordinate conversion and linear dimming driving. In addition, the system also adopts a chromaticity stability feedback and user preference learning mechanism to continuously optimize the dimming and toning effects. The problems of safety problem, light color drift and lack of intelligent adjustment capability caused by stroboflash of digital dimming in the prior art are solved. By introducing a linear dimming mode and applying an AI fuzzy control algorithm, the system can perform dynamic adjustment according to environmental changes and automatically select the most suitable dimming and toning strategy, thereby ensuring that accurate and stable color temperature and brightness are output in different environments.
This invention discloses a hardware autonomous control method and system based on spatial intelligence and self-evolutionary learning, an electronic device, and a computer-readable storage medium. The system includes: a multi-source device discovery module for automatically scanning intelligent hardware devices and establishing a unified device model through multiple communication protocols; a capability reflection module for automatically extracting device control capabilities and parameter constraints from protocol metadata; a skill management module for storing and loading skill packages and establishing a mapping index from device type to skill package; a rule engine module for performing millisecond-level deterministic evaluation of sensor data and generating control commands; an intelligent decision engine module for orchestrating planning nodes and execution nodes based on state diagrams and making context-aware control decisions through a large language model; a command security module; an execution verification module; a multi-layer memory module; a preference learning module; and a multi-layer self-evolutionary engine module.
This invention discloses a method, system, device, and storage medium for intelligent chassis control driven by occupant personalized preferences based on a vision-language-action model, belonging to the field of intelligent vehicle control and intelligent chassis collaborative control technology. The method includes: collecting multi-source information; constructing a personalized dynamic experience preference profile of the occupant; performing scene-preference multimodal joint modeling to form a unified scene-preference semantic representation; inputting the vision-language-action model to generate personalized chassis style intent; generating dynamic performance constraints and controller parameter adjustment targets; generating specific chassis control targets; executing multi-actuator personalized collaborative control; and performing parameter recovery and preference learning updates. This method realizes vision-language-action joint reasoning between occupant personalized preferences, visual scene semantics, and chassis control actions, improving the overall performance of autonomous and assisted driving vehicles in terms of comfort, stability, safety, and personalized experience.
The system according to the embodiment aims to simplify the process of travel planning and to make personalized suggestions according to the user's preferences.SOLUTION: A system according to an embodiment includes a user preference learning unit, a personalization proposal unit, a real-time information acquisition unit, an automatic budget control unit, a collective reservation unit, and a customization unit. The user preference learning unit learns the user's preference. The personalized suggestion unit suggests tourist spots, restaurants, and activities personalized based on the user's preference. The real-time information acquisition unit acquires real-time information of a travel destination. The automated budget manager optimizes the travel plan based on the user's budget. The collective reservation unit makes a collective reservation based on the travel plan. The customization unit customizes the travel plan in accordance with a user's request.SELECTED DRAWING: Figure 1
The invention relates to the technical field of building control, and discloses a building facade control method and system based on network fine tuning, a controller and a storage medium, and the method comprises the steps: carrying out the standardizationprocessing of real-time environment data, and obtaining an environment state vector; inputting the environment state vector into an initial controller, controlling an external facade component and receiving feedback data of a target user; constructing a preference pair in the current period according to the feedback data so as to optimize the initial controller to obtain a target controller; and inputting the environment state vector of the target area in the next period into a target controller so as to adjust the external facade component. According to the method, through deep reinforcement learning driven by deep fusion data and a lightweight online preference learning technology, efficient management of external facade components is improved, and the adaptive capacity, the individuation degree and the multi-target real-time optimization efficiency of equipment are improved.
The invention provides a tool integrated reasoning method through self-evolution preference learning. Comprising two parts: 1, training data construction based on information entropy; 2, a multi-stage self-evolution training normal form; the training data construction process based on the information entropy is used for constructing training data and comprises a data source screening process and an entropy guide sampling process; the multi-stage self-evolution training normal form is used for improving the tool integration reasoning capability of the model, and comprises a two-stage normal form of supervision fine tuning and self-evolution direct preference optimization; according to the reasoning method, for a given input instruction I and a large language model theta, a final answer y is generated through a calculation process. By means of the technical scheme, the effects that efficiency and accuracy are improved, the reasoning process is more stable and simpler, the training data construction cost is lower, and the model generalization ability is higher can be achieved.
This invention discloses an AI customer service interaction strategy adaptive adjustment method based on reinforcement learning, comprising the following steps: acquiring multi-source interaction data and constructing a state-action trajectory input structure; combining human feedback on customer service interaction quality preferences to construct trajectory preference comparison samples and preference condition vectors; constructing a human preference learning framework; constructing a conditional preference reward generation structure within the human preference learning framework to generate trajectory preference scores and preference uncertainty representations; constructing a preference attribution credit allocation structure to generate dense reward sequences corresponding to interaction rounds; performing adaptive preference updates based on online interaction data; constructing a reinforcement learning strategy optimization process to generate an adaptive strategy model; and using the adaptive strategy model to execute online interactions and iteratively update the learning process. This invention achieves fine-grained optimization, stable updates, and continuous adaptive adjustment of customer service strategies.
The invention discloses a man-machine cooperation method for solving human deviation. The man-machine cooperation method comprises the steps that initialization is carried out; iteratively executing batch Thompson sampling, batch Thompson sampling, preference query and data updating and Gaussian process posteriori updating until the maximum number of iterations is reached; and after iteration is finished, returning an optimal action corresponding to the maximum potential function value in the action space. According to the embodiment, the long-tail preference relationship problem is fundamentally solved. In the batch Thompson sampling stage, diversified candidate preference pairs are generated through an adaptive covariance scaling factor and a double-independent sampling mechanism; in the sub-mode marginal gain evaluation stage, the dominant effect of head preference is effectively inhibited by utilizing the profit decreasing characteristic and marginal gain maximization of a sub-mode function; and in a Gaussian process posteriori updating stage, the observed preferences are integrated into a Bayesian framework, and posteriori distribution is refined to guide subsequent sampling. By optimizing the preference learning process, the preference query times are remarkably reduced, and the learning efficiency and accuracy are improved.
An object of a system according to an embodiment is to provide information personalized on the basis of an action history or hobbies and preferences of a user in real time.SOLUTION: A system includes a collection unit, a learning unit, a generation unit, and a provision unit. The collection unit collects an action history or hobbies and preferences of a user. The learning unit learns the data collected by the collection unit. The generation unit generates personalized information on the basis of the data learned by the learning unit. The provision unit provides the information generated by the generation unit to the user in real time.SELECTED DRAWING: Figure 1
The invention discloses a wood texture aesthetic feature extraction and personalized matching system, which belongs to the technical field of wood texture analysis and calculation aesthetics and comprises an image acquisition module, an aesthetic feature extraction module, a preference learning module, a similarity calculation module, a personalized recommendation module and a texture continuity optimization module. The method comprises the following steps: quantitatively extracting multi-dimensional aesthetic characteristic parameters of wood through multi-scale characteristic analysis and visual perception, constructing a personalized preference model based on a user selection history, calculating weighted similarity between to-be-matched wood and user preference, generating a personalized recommendation list, and providing an optimal arrangement scheme for splicing multiple pieces of wood. The technical problems that in the prior art, wood texture aesthetic feature quantification is not comprehensive, a personalized recommendation mechanism is lacked, and the multi-piece splicing effect is poor are solved, and the method is suitable for the fields of high-end furniture manufacturing, interior design, wood product personalized customization and the like.
The invention relates to the field of network communication, and provides a network strategy optimization method and system based on user behaviors and preference learning. The method comprises the following steps: collecting strategy operation behaviors and network quality sensing data of a user to obtain a user behavior data set; inputting the user behavior data set into a machine learning model for training so as to learn an association relationship between user behavior characteristics and network quality indexes, and obtaining a user preference weight vector; obtaining a quality parameter and available network strategies of the current network environment, and performing quantitative evaluation on the plurality of available network strategies according to the user preference weight vector to obtain a comprehensive score of the plurality of available network strategies; and performing priority ranking on the available network strategies based on the comprehensive score, and selecting a target network strategy according to a strategy availability state and a score result. According to the method, the personalized preference of the user on the network quality index can be learned, intelligent evaluation and dynamic switching of the network strategy are realized, and the user experience is improved.
The invention discloses a federal multi-modalpreference learning and alignment method for a mode lacking condition, and aims to solve the problem of unstable alignment caused by image-text data mode lacking, preference noise and cross-domain conflict in multi-clientprivacy protection. According to the method, a modal state is identified and a conditional preference pair is constructed at a client: cross-modal consistency scoring is carried out to screen good and bad responses when images and texts are complete, a visual attribute lexicon is introduced to punishment suppression assume when the texts are lack of texts, and entity existence verification punishment illusion is carried out when the images are lack of texts; noise updating is reduced in combination with sample-level and client-level two-level gating based on prediction uncertainty; and the server side forms multi-preference experts by utilizing low-dimensional preference signature clustering and issues and fuses the multi-preference experts according to similarity routing. According to the method, the consistency and robustness of multi-modal generation are improved in a modal-lacking scene, and negative migration and illusion are reduced.
The application provides an interpretable reward model construction method based on a sparse self-encoder, and the method comprises the following steps: constructing and pre-training a sequence-level sparse self-encoder, and constructing an interpretable reward model, wherein the encoder part of the pre-trained sparse self-encoder is integrated after the first layer of a basic language model, and all original layers after the first layer of the basic language model are discarded; the reward model is trained based on preference data, wherein the encoder parameters of the sparse self-encoder are frozen, the parameters of the front layers of the basic language model and the weights of a linear value head are trained, the weights of the linear value head are optimized by using a pair of preference data through a preference learningloss function, and the model learns the contribution weight of each interpretable feature to human preference. The method can be simply and efficiently applied to various real scenes, and is particularly suitable for reinforcement learning training of a strategy model under dynamic preference and large-scale data filtering.
The application discloses an ultrasonic tongue silent speech recognition method based on modal transfer learning, relates to the technical field of speech information processing, and proposes a multi-stage and multi-modal modeling pipeline for two main challenges of SSR, namely, a limited data set and a confusing silent signal. Four-step processes are introduced to avoid overfitting and two additional modules are used to extract multi-modal information. Based on a joint training strategy, modal preference learning (MPL) is proposed to promote the optimization of the target task by utilizing cross-modal prior knowledge.
The invention belongs to the technical field of cross-domain sequence recommendation, and discloses a cross-domain sequence recommendation method based on causal inference and preference evolution. Through a cue word template designed based on cross-domain co-occurrence frequency, article semantic information is enhanced by using a large language model. Performing intra-domain sequence preference learning through the domain-specific interest evolution model to obtain intra-domain preferences; and performing cross-domain sequence preference learning through a time sequence-domain dual-condition hybrid expert mechanism to obtain cross-domain preference. Causal depolarization is designed to decouple the user intra-domain preference and the user cross-domain preference from the user intra-domain activeness and the user cross-domain activeness, recommendation is generated, and the accuracy of a cross-domain recommendation result is improved. According to the method, by means of a causal enhanced preference learning mechanism, the limitation of an existing method is effectively solved, and the most advanced performance is achieved in single-domain and cross-domain sequence recommendation benchmark tests.
The invention relates to the technical field of automobile travel service, discloses a travel comprehensive cost optimization method, system and equipment based on a multi-modallarge model, and aims to solve the problems of poor travel experience and high comprehensive cost caused by difficulty in considering charging site selection, accommodation hotel screening and travel cost management and control during long-distance travel. The method comprises the steps of collecting travel associated data based on travel demand information; constructing a comprehensive cost function including charging cost, hotel cost and travel time cost, and defining a weight coefficient based on user preference learning; calculating charging cost, hotel cost and travel time cost in the comprehensive cost function according to travel association data; and based on a calculation result of the comprehensive cost function, screening, sorting and optimizing the multiple groups of charging-hotel-path combination schemes by adopting a decision inference engine, generating an optimal electric travel scheme, and outputting the optimal electric travel scheme to a vehicle-mounted terminal for a vehicle owner to select, so that the comprehensive cost optimization and travel experience improvement of the long-distance travel of the vehicle are realized.
The embodiment of the invention provides a model training and content generation method, computing equipment, a storage medium and a product. The method comprises the steps that multiple pieces of first sample popularization content generated based on sample input information are obtained, and the sample input information comprises object information of first sample objects; positive sample content meeting the preference requirement and negative sample content not meeting the preference requirement are determined from the multiple pieces of first sample promotion content; based on the sample input information, the positive sample content and the negative sample content, performing preference learning on the target generation model to train the target generation model; the trained target generation model can generate target promotion content based on the target input information; the target input information comprises object information of the target object. According to the technical scheme, the target generation model is trained in a preference learning mode, the quality stability of the target promotion content generated by the target generation model is improved, and then the conversion rate of the target object corresponding to the promotion content is improved.
The present application belongs to the technical field of multi-modallarge model alignment and preference learning, and particularly relates to a multi-modal preference optimization method and system based on bidirectional distribution alignment. The method generates candidate response sets corresponding to the original image and the edited image respectively based on the original image, the edited image and the question under the current model parameter distribution, then for each candidate response, positive samples are screened based on the confidence score and the reference consistency score with the reference response, and negative samples are screened through an entropy-guided negative sample mining strategy; subsequently, the positive and negative samples corresponding to the original image and the edited image are input into the model, the overall loss of the model is calculated based on the image-level and response-level contrastive loss, and the model parameter distribution is updated until the model converges. The present application significantly reduces the hallucination rate through bidirectional alignment of the model parameter distribution and the data distribution, and introduces an entropy-guided negative sample mining strategy in the direction of data adaptation model, thereby realizing targeted hallucination suppression.
This application relates to the field of medical information extraction technology, and particularly to a biomedical information extraction method based on two-stage fine-tuning and preference optimization. The method includes: optimizing a base model using a two-stage fine-tuning and preference optimization strategy, and then using the optimized model to extract biomedical information. This application designs a progressive preference learning framework, employing an improved DPO-Positive algorithm and post-fine-tuning to enhance model accuracy for medical information extraction tasks; it also designs a multi-dimensional preference dataset and its automated generation method to reduce the burden of manual annotation and improve model fault tolerance and generalization ability; and combines efficient parameter fine-tuning and dual-model correctionverification to achieve high-performance, standardized output with limited computing power, improving knowledge extraction efficiency and reliability.
The invention discloses a personalized drug recommendation method based on multi-target deep reinforcement learning, and belongs to the field of bioinformatics. The method comprises the following steps: firstly, acquiring cell line multi-omics and drug sequence data, and constructing a multi-source feature extraction and projection module to extract potential feature representation; constructing a dynamic historical preference learning and deep interactive fusion module, and performing deep fusion on multi-source features and time sequence preferences by using a gating residual mechanism to generate comprehensive state representation; an intelligent agent decision module is constructed, a drug recommendation strategy and comparison feature representation are generated by utilizing a Hybrid-Actor network double branch, and a curative effect value and a safety value are evaluated by utilizing a Dual-Critic network; and constructing a joint objective function in combination with self-supervised comparative learning and multi-objective reward, performing iterative optimization on the model, and finally outputting a high-efficiency and low-toxicity personalized drug recommendation list. According to the method, the toxicity risk can be remarkably reduced while the high curative effect of the recommended medicine is ensured, and a scientific basis is provided for clinical personalized medicine recommendation.
The application relates to the technical field of model training, and particularly discloses a security foundation big model based on preference reinforcement learning and a training method, which comprises the following steps: a processing module is used to obtain all positive samples and all negative samples of an input video group based on the portrait description of all videos of the input video group; a training module is used to obtain all comparison sample pairs based on all positive samples, all negative samples and a preset unlabeled video group of the input video group; a calculation module is used to obtain the iterative sorting of all comparison sample pairs based on all comparison sample pairs and a preset preference learning technology; and an iteration module is used to obtain the iterative screening straight line of the security foundation big model based on the iterative sorting of all comparison sample pairs, so that the security foundation big model is obtained. The application can train a model which can accurately deduce the correct portrait description of an input video by using a small amount of training data.
An object of a system according to an embodiment is to learn preferences and needs of a user and actively propose an idea based on the learned preferences and needs.SOLUTION: A system includes a preference learning part, an idea proposing part, and a generation AIBuilder part. The preference learning unit automatically learns preferences and needs of the user through interaction with the user. The idea proposing section actively proposes an idea based on the user's preferences and needs learned by the preference learning section. The generation AIBuilder component provides functionality for users to build their own generation AI.SELECTED DRAWING: Figure 1
The invention discloses a self-adaptive closed-loop sleep nerve regulation and control system based on implicit preference learning. The self-adaptive closed-loop sleep nerve regulation and control system comprises a multi-modalsignal acquisition module, a nerve state real-time monitoring module, an implicit preference driving decision module, a multi-modal stimulation execution module and a self-evolution feedback learning engine, and the multi-modalsignal acquisition module is used for noninvasively acquiring physiological and environmental data of a user and sending the acquired data stream to the neural state real-time monitoring module and the self-evolution feedback learning engine. According to the invention, by constructing a full-closed-loop neural feedback architecture of perception-decision-regulation-learning, the crossing from passive monitoring to active optimization is realized, and the purpose is to solve the core problems in the prior art that the sleep intervention means are lack of individuation, the feedback mechanism is rigid, and the user experience is disjointed from the physiological benefit.