Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
84 results about "Preference learning" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
Preference learning is a subfield in machine learning, which is a classification method based on observed preference information . In the view of supervised learning, preference learning trains on a set of items which have preferences toward labels or other items and predicts the preferences for all items.
The invention discloses a machinebehavior learning method based on emotion driving and human expert feedback, and the method comprises the following steps: a) fusing environment image features extracted by BLIP-2 and text instruction semantic features analyzed by GPT-4, and forming cross-modal input; b) training a basic VLA model through supervised fine tuning (SFT) by using human expert remote control trajectory data and cross-modal input to obtain a basic behavior strategy; c) combining an emotion recognition module with a multi-head self-attention mechanism, fusing emotional interaction dependency into a basic strategy, and generating high-order emotion driven behavior representation; and d) inputting the high-order emotional behavior representation into a reinforcement learning module, storing a track by using a Replay Buffer, carrying out optimization through human expert feedback preference learning, and outputting a final behavior strategy. Compared with an existing method, the method has the advantages of being high in multi-modalfeature extraction capacity, high in emotion fusion degree, sufficient in expert feedback utilization and the like, and the response accuracy and interaction experience of a machine to human instructions and emotions can be improved to a certain degree.
The invention relates to the technical field of artificial intelligence, in particular to an intelligent content evaluation and optimization method and system based on multi-standard preference learning, and the method comprises the steps: constructing target user preference data, and generating an evaluation track containing evaluation rules and judgment results; based on the evaluation trajectory, screening samples and distributing credits through sorting and consistency rules to obtain preference pair training data; jointly training a generative reward model by adopting response supervision fine tuning and a direct preference optimization strategy; recombining the original evaluation trajectory into alternate accepting and rejecting samples, forming long thinking chain training data, and further training to obtain a final model; and evaluating and optimizing the alignment degree of the generated text and the user preference through the final model. According to the method, through multi-stage optimization and process supervision, the alignment degree and the overall performance of the language model and human preference are improved, the problems of composite errors, data sparseness and the like of a traditional reward model are solved, and the method is excellent in performance in out-of-distribution evaluation.
The invention relates to the technical field of image processing, and provides an indoor three-dimensional image intelligent rendering method based on image processing. Comprising the following steps of indoor multi-view image acquisition, image preprocessing, intelligent analysis of indoor scene elements, construction of an indoor three-dimensional initial model with attributes, intelligent generation of rendering parameters, adaptive LOD real-time rendering and intelligent interaction optimization. According to the intelligent interaction optimization system, through deep fusion of natural languageprocessing, user preference learning and real-time rendering technologies, a set of efficient, visual and personalized virtual scene rendering interaction process is constructed. According to the method, the problems that traditional graphic software is complex in operation, high in learning cost and low in debugging efficiency are solved, and the user satisfaction and creation efficiency are remarkably improved through an intelligent recommendation and rapid iteration mechanism.
The invention relates to the technical field of intelligent lighting control, in particular to a multi-user sharing-oriented self-adaptive lighting intelligent identification control method, which comprises the following steps of: identifying user identities through a multi-user feature identification system, acquiring historical use habit data, generating personalized lighting parameters based on the data and current environment data, the group preference learningsystem constructs similar user clusters to optimize shared space illumination; the conflict adjustment system identifies and processes multi-user demand conflicts; the priority distribution system dynamically adjusts user control authority; the feedback learning system obtains user feedback, calculates satisfaction, updates a preference model and continuously optimizes an illumination control strategy; intelligent lighting control based on user features and group preferences is achieved, conflict processing and continuous learning capabilities are achieved, rapid and accurate recognition of users is achieved through a face recognition technology, lighting parameters are automatically adjusted in combination with historical use habit data, and personalized lighting experience is provided; dynamic changes of user demands are captured, and the accuracy of illumination scene recommendation is improved.
The invention relates to an intelligent clothing recommendation system based on image recognition and a recommendation method thereof, and the system obtains a human body image through a multi-modal data collection module, constructs a three-dimensional body model in combination with a deep learningalgorithm, and achieves the precise measurement of key parameters such as shoulder breadth and chest circumference. And customized recommendation schemes are generated based on body types such as apple types and pear types. The clothing recommendation module integrates fashion trend analysis and a scene adaptationalgorithm, provides style screening, single article combination and dynamic AR fitting functions, and supports 360-degree rotation observation and fabric texture simulation. The virtual fitting module generates dynamic wrinkles of clothes through motion capture, and the error is controlled within 150ms. The system integrates an e-commerce platform and a social sharing function, supports user preference learning and intelligent reminding, improves the accuracy of body type measurement, and improves the recommendation matching degree. Through three-dimensional body shape analysis, dynamic physical simulation and multi-mode interaction, the garment fitting precision and the user experience are remarkably improved.
The invention discloses an intelligent exhibition item module clustering management method and system supporting multi-scene switching. The intelligent exhibition item module clustering management method comprises the following steps: step 1, real-time state acquisition; 2, generating a dynamic strategy; step 3, content updating and issuing; step 4, operation monitoring and fault processing; 5, multi-scene adaptation is carried out; and step 6, audience preference learning. According to the method, intelligent collaborative scheduling of exhibition item clusters is realized through a real-time state acquisition and dynamic strategy generation mechanism, the system can automatically optimize resource allocation based on multi-dimensional load evaluation and an audience interest model, the problem of resource scrambling caused by independent operation of traditional exhibition items is solved, content hot update and cluster-level synchronization are supported, and the system is suitable for popularization and application. The bottleneck of traditional display content homogenization is broken through, meanwhile, a predictive operation and maintenance closed loop is constructed by a dual-channel redundancy monitoring and three-level fault response system, the equipment downtime risk is remarkably reduced, and the system stability is improved.
The invention discloses a layering and preference learning-based personalized dynamic treatment strategy generation system for Parkinson's disease, which comprises a dynamic state and preference representation module, a preference-based reward function learning module, a layering treatment strategy generation module and a treatment path simulation and interactive output module, the system learns preferences from doctor-patient interaction and generates layered, dynamic and explainable personalized long-term treatment strategies, so that making of personalized long-term treatment decisions is achieved, and the problems that in an existing Parkinson's diseasereinforcement learning scheme, an award function is fixed, and long-term planning ability is lacked are solved.
The invention belongs to the field of artificial intelligence, particularly relates to an intelligent sound fieldadaptive system of digital professional sound equipment, and aims to solve the problems of inaccurate sound field regulation and control, response lag, dependence on manual tuning and the like in a complex acoustic environment. The system comprises a sound field sensing module, an acoustic modeling and analysis module, a self-adaptive sound field regulation and control engine, a multi-channel digital audioprocessing unit and a feedback optimization module, and dynamic sound field modeling, real-time audio processing and environment self-adaptive regulation and control are realized through distributed sensing, hybrid modeling, multi-target optimization and closed-loop feedback. The system supports rapid re-calibration, multi-scene memory and user preference learning, ensures voice clarity and music fidelity, balances full-field hearing consistency, and significantly reduces manual intervention requirements.
The invention provides a text classification model optimization method and a text classification method and device.The method comprises the steps that each text in an original data set is divided, and semantic units of the texts at different levels are obtained; respectively enhancing the semantic units of each level by adopting different data enhancement modes, obtaining a preset number of supplementary texts for each text, and arranging the supplementary texts into an original data set to obtain an enhanced data set; constructing a preference reward function based on the keyword semantic score and the global semantic score; and on the enhanced data set, performing fine adjustment on the original text classification model by using a preference reward function to obtain an optimized text classification model. According to the method, a multi-strategy fine-grained data enhancement method is realized, a preference award function is constructed, and preference learning is introduced, so that model training is guided in the process of optimizing a text classification model, and the text classification model with higher accuracy and robustness is obtained through optimization.
The invention discloses a video rapid editing method based on artificial intelligence, and particularly relates to the technical field of video processing. Comprising the steps of S01, video data feature intelligent acquisition, S02, video data feature intelligent identification, S03, video clip intelligent generation, S04, video clip effect intelligent management, S05, video clip strategy evaluation and S06, user preference learning. A video frame to be edited is converted into a comparable standard numerical sequence from multi-dimensional quantitative evaluation, the intelligence and efficiency of editing are improved, the editing fluency of a video frame with a poor editing result is improved by continuously iterating and optimizing a video editing strategy through an editing effect fluency evaluation coefficient between adjacent video frames, and the editing efficiency is improved. By learning the preference of the user, the generated initial editing sequence conforms to the habits of the user.
The invention relates to the technical field of network data security, discloses a fact checking method and system based on multiple agents, a storage medium and electronic equipment, and aims at solving the problem that the understanding ability of complex declarations is insufficient in the prior art. Decomposing the complex original declaration into a plurality of sub-declarations which cannot be subdivided; in order to solve the problem that external knowledge is lacked to assist in verifying facts in the prior art, a retrieval agent taking Qwen2.5-72B as a basic model is constructed, and fact verification evidences are retrieved from a fact verificationknowledge base or the Internet for sub-declarations; in order to solve the problem that in the prior art, a fact checking result lacking logic verification is directly output, Qwen2.5-72B trained through a human preference learningalgorithm is constructed to serve as a decision-making agent, whether all evidences ei and answers ai in a sub-declaration sequence logically support an original declaration or not is evaluated, and a final fact checking conclusion is output.
Embodiments described herein provide for training an artificial intelligence model to become a preference-aware model. The artificial intelligence model preferences as the artificial intelligence model trains. Reinforcement learning is used to train experts in the artificial intelligence model such that each expert is trained to converge to a unique preference. The architecture of the artificial intelligence model is highly flexible. Upon executing a trained model, users can select automatically images according to various preferences based on medical professional preferences, geographic preferences, patient anatomy, and institutional guidelines.
The invention relates to the technical field of intelligent sunshade equipment control, in particular to an automatic controlsystem and method for a sunshade, and aims to solve the problem that an existing automatic controlsystem for the sunshade generally does not fully integrate multi-dimensional environment parameters and user intervention behaviors in the aspect of strategy generation, and the strategy generation efficiency is high. The problems that energy consumption, comfort and illumination uniformity are difficult to optimize synchronously under a dynamic meteorological condition, so that a control strategy is lack of individuation and foresight are solved; multi-dimensional environment semantic modeling, short-term illumination trend prediction and user preference learning are fused through a sun shading strategy dynamic generation module, a dynamic weighted multi-target optimization mechanism is constructed, energy consumption, comfort and illumination uniformity are synchronously considered under the condition that equipment constraints are met, a control target and an evaluation standard are updated in real time based on user intervention records, and the system is high in practicability. Accurate, prospective and highly personalized sunshade strategy generation is realized, and the intelligent level and comprehensive performance of the system are remarkably improved.
The invention discloses a memory enhancement-based large language model lifelong alignment method. The method comprises the steps of focus preference optimization and short-time to long-time memory consolidation. The focus preference optimization adaptively focuses the learning focus on a new preference sample or a preference sample with an uncertain model through an improved preference learningloss function, meanwhile, the updating amplitude of fully learned knowledge is reduced, and a historical alignment result is protected while new preference is learned; the short-term to long-term memory consolidation is used for simulating a human memory mechanism, denoising is performed on short-term parameter update through singular valuedecomposition, an update part conflicting with past knowledge is identified and suppressed by projecting to a historical knowledge subspace, and finally refined conflict-free knowledge is integrated into long-term parameters of a model. According to the method, knowledge of the model is effectively accumulated and reserved in continuous and diversified alignment tasks, catastrophic forgetting is remarkably inhibited, and the stability, reliability and alignment consistency of the model in a dynamic and long-term deployment environment are improved.
The invention relates to the technical field of model training, and particularly discloses a security basic large model based on preference reinforcement learning and a training method, and the model comprises a processing module which is used for obtaining all positive samples and all negative samples of an input video group based on the portrait description of all videos of the input video group; the training module is used for obtaining all comparison sample pairs based on all positive samples and all negative samples of the input video group and a preset unlabeled video group; the calculation module is used for obtaining an iterative sequence of all the comparison sample pairs based on all the comparison sample pairs and a preset preference learning technology; and the iteration module is used for obtaining an iteration screening straight line of the security and protection basic large model based on the iteration sorting of all the comparison sample pairs so as to obtain the security and protection basic large model. According to the invention, a model which can be trained by using a small amount of training data and can accurately derive correct human portrait description of an input video is realized.
The application discloses a kind of based on user text generation content's minority preference learning method in the field of information retrieval, comprising the following steps: data preprocessing operation is carried out to the user text generation content obtained;The data obtained by pre-processing establishes a hierarchical Bayesian model, obtains joint distribution model;Model parameters are learned by Gibbs sampling method, and the formula of mass preference distribution and minority preference distribution is obtained;The meaning of the user minority preference based on user text generation content is analyzed using the learned model parameters;The target user under the minority preference is found using the user minority preference distribution.The method of the application distinguishes mass preference and minority preference from the perspective of user preference, identifies the specific meaning of user minority preference using the good interpretability of hierarchical Bayesian method, provides the opportunity for small and medium-sized enterprises to enter suitable niche market, and the minority preference distribution of each user is beneficial to the enterprise to find out the target user of relevant niche market.
The invention discloses a linear dimming type mixed light illumination system and method based on AI fuzzy control, and belongs to the technical field of LED illumination, the system collects information such as ambient brightness, time, human body activity and the like in real time through an environment sensing module, processes data through an AI-driven fuzzy control algorithm, and outputs target color temperature and brightness level through reasoning. And then, accurate mixed light output is realized through color temperature and brightness coordinate conversion and linear dimming driving. In addition, the system also adopts a chromaticity stability feedback and user preference learning mechanism to continuously optimize the dimming and toning effects. The problems of safety problem, light color drift and lack of intelligent adjustment capability caused by stroboflash of digital dimming in the prior art are solved. By introducing a linear dimming mode and applying an AI fuzzy control algorithm, the system can perform dynamic adjustment according to environmental changes and automatically select the most suitable dimming and toning strategy, thereby ensuring that accurate and stable color temperature and brightness are output in different environments.
The present application patent belongs to the field of big data analysis technology and personalized recommendation system, and particularly relates to an online learning behavior personalized recommendation system based on clustering analysis method, which mainly comprises four functional modules, namely an online learning behavior data information management module, a data analysis modeling module, a student learning portrait display module and a same-type learning friend personalized recommendation module. The system firstly collects and cleans data from an online learning platform, constructs an online learning student portrait label system, and based on the clustering analysis method, creates an online student portrait for each online student from three dimensions of learning attitude style, learning interest preference and learning level ability. Meanwhile, based on the same-type learning friend personalized recommendation model, students with similar online learning behaviors are recommended to each other to further stimulate the learning enthusiasm of students through mutual exchange and discussion.
The invention discloses a preference learning-based code repair system and method. A code repair enhancement module of the system trains the large language model based on the code data set so as to construct an initial code repairer; a repair preference data generation module generates candidate repair codes from the programming tasks and the error codes through an initial code repairer, and constructs a preference learningdata set after preference data selection; the code preference learning module is used for training the initial code restorer through the preference learning data set so as to obtain a final code restorer; and the code repair generation module generates a repair code with a small modification range for the to-be-repaired code data through the final code repairer. According to the method and the device, the repair code which is accurately repaired and has a smaller modification range can be intelligently generated according to actual requirements in a development scene, so that code problems are quickly positioned and solved, unnecessary modification of the code is effectively reduced, the consistency and readability of the code are kept, the time cost of manual debugging is reduced, and the development efficiency is improved.
This invention discloses a hardware autonomous control method and system based on spatial intelligence and self-evolutionary learning, an electronic device, and a computer-readable storage medium. The system includes: a multi-source device discovery module for automatically scanning intelligent hardware devices and establishing a unified device model through multiple communication protocols; a capability reflection module for automatically extracting device control capabilities and parameter constraints from protocol metadata; a skill management module for storing and loading skill packages and establishing a mapping index from device type to skill package; a rule engine module for performing millisecond-level deterministic evaluation of sensor data and generating control commands; an intelligent decision engine module for orchestrating planning nodes and execution nodes based on state diagrams and making context-aware control decisions through a large language model; a command security module; an execution verification module; a multi-layer memory module; a preference learning module; and a multi-layer self-evolutionary engine module.
The invention relates to the technical field of smart city equipment control, and discloses a smart city equipment closed-loop control method and system based on machine learning, and the method comprises the steps: obtaining a user advanced instruction and implicit preference through a natural languageprocessing and behavior pattern analysis technology, establishing multi-tenant demand mapping; analyzing the tenant demand data by using a Bayesian preference learningalgorithm, and generating a multi-dimensional tenant portrait; seeking an optimal balance point between the tenant satisfaction and the energy efficiency, and generating a control decision scheme; implementing a dynamic space-time partition control strategy; when a system component fails or an unforeseen scene is encountered, a replacement control program is automatically synthesized, and system self-repairing is achieved; the multi-tenant satisfaction degree is improved, and the user satisfaction degree mean value is improved; the adaptive capacity of the system is enhanced, the abnormal scene processing proportion without manual intervention is improved, and the recovery time after the system fails is shortened.
The invention relates to a context preference learning method, device and equipment based on a large language model. The method comprises the following steps: configuring a task environment and defining performance indexes; automatically generating multiple groups of initial reward functions through a large language model; training a plurality of reinforcement learning agents in parallel and collecting behavior data; automatically identifying an optimal reward function and a worst reward function based on a weighted scoring mechanism; and driving a large language model to generate an improved function in combination with comparison difference information, and performing iterative optimization. According to the scheme, the bottleneck of manual design can be broken through, and dynamic weight adjustment and cross-scene generalization of the reward function are realized; the decision-making efficiency and safety are obviously improved in the scenes of customer service, industrial control and the like; the labor cost is reduced through unsupervised preference evaluation, and continuous adaptive optimization of the strategy model is supported.
The invention provides an information security dimension-based model training optimization method and device, and the method comprises the steps: automatically generating security training data and risk training data needed for training a first multi-modallarge model through a first large model according to multi-modal sample data comprising security sample data and risk sample data; furthermore, on the basis of the security training data and the risk training data, a preference learning mode is adopted to carry out iterative training on the first multi-modallarge model to obtain a second multi-modal large model with a better copywriting generation capability, and the copywriting content is generated by the second multi-modal large model, so that security and compliance requirements are met.
This invention discloses a method, system, device, and storage medium for intelligent chassis control driven by occupant personalized preferences based on a vision-language-action model, belonging to the field of intelligent vehicle control and intelligent chassis collaborative control technology. The method includes: collecting multi-source information; constructing a personalized dynamic experience preference profile of the occupant; performing scene-preference multimodal joint modeling to form a unified scene-preference semantic representation; inputting the vision-language-action model to generate personalized chassis style intent; generating dynamic performance constraints and controller parameter adjustment targets; generating specific chassis control targets; executing multi-actuator personalized collaborative control; and performing parameter recovery and preference learning updates. This method realizes vision-language-action joint reasoning between occupant personalized preferences, visual scene semantics, and chassis control actions, improving the overall performance of autonomous and assisted driving vehicles in terms of comfort, stability, safety, and personalized experience.
The system according to the embodiment aims to simplify the process of travel planning and to make personalized suggestions according to the user's preferences.SOLUTION: A system according to an embodiment includes a user preference learning unit, a personalization proposal unit, a real-time information acquisition unit, an automatic budget control unit, a collective reservation unit, and a customization unit. The user preference learning unit learns the user's preference. The personalized suggestion unit suggests tourist spots, restaurants, and activities personalized based on the user's preference. The real-time information acquisition unit acquires real-time information of a travel destination. The automated budget manager optimizes the travel plan based on the user's budget. The collective reservation unit makes a collective reservation based on the travel plan. The customization unit customizes the travel plan in accordance with a user's request.SELECTED DRAWING: Figure 1