Method and device for interactive controlling on adult product based on multimodal model, and adult product
Patent Information
- Application Number
- US19/312858
- Authority / Receiving Office
- US · United States
- Patent Type
- Patents(United States)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Currently, an amplitude or frequency with which the adult product works may be adjusted manually or remotely during using the product, but cannot be adjusted according to sexual activities by users and thus the user may feel less interactions so that the user experience in using such products may be degraded.
[0089]Compared to the prior art, the embodiments of the present disclosure have at least the following advantageous effects.
Smart Images

Figure US12746176-D00000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is a Continuation of U.S. patent application Ser. No. 19 / 275,239 filed Jul. 21, 2025, the entire contents of which are incorporated herein by reference.TECHNICAL FIELD
[0002] The present disclosure relates to the field of interactive controlling on adult products based on artificial intelligence, specifically to a method and device for interactive controlling on adult product based on a multimodal model, and adult product.BACKGROUND
[0003] With the advancement of social civilization, more and more attentions are made on addressing on physiological needs appropriately. Adult products, also known as sex toys, are widely used. Examples of adult product may include vibrators, massager, artificial vaginas, and silicone dolls. These products are primarily used to enhance or assist sexual experiences for adults to facilitate self-stimulation or use with a partner for sexual pleasure.
[0004] Currently, an amplitude or frequency with which the adult product works may be adjusted manually or remotely during using the product, but cannot be adjusted according to sexual activities by users and thus the user may feel less interactions so that the user experience in using such products may be degraded.BRIEF SUMMARY
[0005] The present disclosure intends to address the problems in the prior art, such as less intelligent interaction, inaccurate user intent recognition, and insufficient personalized experience in adult products, by providing a method and a device for interactive controlling on adult product based on a multimodal model, and an adult product. In the present disclosure, intelligent interaction may be achieved between adult products and users with the combination of multimodal information processing and deep learning models, so as to improve the user experience significantly.
[0006] In a first aspect of the present disclosure, a method for dynamically controlling physical actuator of adult product, including:
[0007] acquiring perception data of a user interacting with the adult product by at least one microphone and at least one sensor, wherein the perception data of the user includes one or more modal data of voice data and tactile data; the sensor includes at least one pressure sensor;
[0008] processing the perception data of the user by using a processor to running a multimodal neural network model, which includes: conducting feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto; applying an attention mechanism to each modal feature vector to obtain modal weights by calculation and generate a weighted cross-modal feature representation that quantifies an emotional state of the user; mapping the weighted cross-modal feature representation to a specific emotion class of the user selected from a predefined set of emotion classes;
[0009] conducting search in a data structure of action parameter vectors to generate a multi-component action parameter vector, A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the specific emotion class of the user; each component of the action parameter vector under searching represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions including at least vibration frequency and contraction intensity; and
[0010] converting the numerical control values of the action parameter vector under searching into drive signals and transmitting the drive signals to the physical actuator, so that the adult product performs physical actions defined by the numerical control values.
[0011] Alternatively, the control instruction further includes one or more of the following: voice control instruction, configured to make selection among preset voice feedback templates and adjust voice parameters according to the emotion class.
[0012] atmosphere control instruction, configured to change one or more of brightness of a light source, color of a light source, flicker frequency of the light source, background music, based on the emotion class of the user.
[0013] Alternatively, the method for dynamically controlling physical actuator of adult product further includes:
[0014] acquiring real-time feedback data from the user, wherein the feedback data includes behavior response; and
[0015] dynamically adjusting predefined action parameter vectors stored in the data structure based on the real-time feedback data.
[0016] Alternatively, the method for dynamically controlling physical actuator of adult product further includes:
[0017] determining a priority of each control instruction based on a preset weight of the user's emotion class;
[0018] adjusting an execution sequence of a plurality of control instructions to ensure collaboration of the instructions and prevent conflicts therebetween;
[0019] acquiring real-time feedback data from the user, including behavioral responses, and adjusting control strategies dynamically according to the real-time feedback data to optimize the user experience.
[0020] Alternatively, the step of conducting feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto and conducting emotion recognition thereon includes:
[0021] converting voice data using a pre-trained language model into voice text and extracting semantic representation vectors, when the perception data includes voice data; the pre-trained language model is subjected to adaptive training in an adult field with respect to vocabulary and expressions in the adult field;
[0022] extracting fundamental audio features from the voice data using Mel-frequency cepstral coefficients; extracting acoustic feature vectors from the fundamental audio features using spectrum analysis, the acoustic feature vectors include fundamental frequency, energy, and harmonic-to-noise ratio; extracting high-dimensional voice representations using a self-supervised pre-trained speech encoder, to determine emotions and tones using an acoustic analyzer;
[0023] processing tactile data by using a multilayer perceptron network to extract tactile feature vectors when the perception data includes tactile data, the tactile feature vectors include one or more of touch pattern, touch force, touch rhythm, touch trajectory, touch area, and action frequency;
[0024] extracting facial expression features based on an improved FaceNet, when the perception data includes visual data, determining gazing duration and gazing frequency using a pupil tracking algorithm to obtain eye contact feature vectors; extracting body posture feature vectors using a human pose estimation model.
[0025] Alternatively, after determining emotion class of the user from a predefined set of emotion classes, the method further includes:
[0026] evaluating confidence of the emotion class by using a neural network.
[0027] Alternatively, the method for dynamically controlling physical actuator of an adult product further includes:
[0028] forming temporal sequence features from the modal features corresponding to each type of modal data within a preset period and determining dynamic changes in emotion recognitions by processing the temporal sequence features using a long short-term memory network, to determine an emotional trend;
[0029] determining a moment of change in emotion, by focusing on transition points in the temporal sequence features with attention weights.
[0030] Alternatively, the method for dynamically controlling physical actuator of an adult product further includes:
[0031] collecting user interaction data and feedback information;
[0032] updating parameters of the multimodal model dynamically using deep learning algorithms and reinforcement learning algorithms;
[0033] adjusting control logic to meet user's personalized needs adaptively and optimizing determining on user intent and control strategies.
[0034] Alternatively, the method for dynamically controlling physical actuator of an adult product further includes:
[0035] performing multi-factor authentication and authorization management on user identity to ensure that only authorized users can control the current adult product.
[0036] encrypting user data in a way of end-to-end to prevent unauthorized access and data leaks.
[0037] monitoring an execution process of the control instructions to block abnormal instructions to ensure safe use.
[0038] Alternatively, the method for dynamically controlling physical actuator of an adult product further includes:
[0039] downloading and installing a latest multimodal model parameters or control logic through a network connection or from a cloud platform, where the update process uses digital signatures and verification mechanisms to ensure the authenticity and integrity of the updated content.
[0040] In a second aspect, the present disclosure further provides a method for interactive controlling on adult product based on a multimodal model, including the following steps:
[0041] acquiring perception data of a user interacting with the adult product, wherein the perception data of the user includes one or more modal data of voice data and tactile data;
[0042] conducting feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto;
[0043] conducting feature-fusing on the extracted modal feature vectors to form cross-modal feature representation so as to determine emotion class of the user from a predefined set of emotion classes;
[0044] conducting search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product; and
[0045] converting the control instruction into drive signals for the adult product, so that the adult product performs actions corresponding to the drive signals.
[0046] Alternatively, the step of the conducting search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product includes:
[0047] conducting search in a data structure of action parameter vectors to generate a multi-component action parameter vector corresponding thereto, A=[a1, a2, . . . , an]. The data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the specific emotion class of the user. The control instructions are configured to adjust one or more of the motion mode, intensity, or frequency of the adult product.
[0048] Each component a1, a2, . . . , an represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions include one or more of vibration frequency, vibration intensity, contraction frequency, contraction intensity, swing amplitude, swing speed, temperature change amplitude, and pressure intensity of tactile feedback. A range of parameters of the physical control dimension corresponds to a preset set of movement patterns.
[0049] Alternatively, the step of conducting feature-fusing on the extracted modal feature vectors to form cross-modal feature representation so as to determine emotion class of the user from a predefined set of emotion classes includes:
[0050] projecting each modal feature vector into a common dimensional feature space by linear mapping, to generate aligned modal features;
[0051] determining attention scores between the aligned modal features by calculation, to determine modal weights for generating a weighted cross-modal feature representation, and generate a weighted cross-modal feature representation that quantifies the emotional state of the user;
[0052] mapping the weighted cross-modal feature representation to a specific emotion class of the user selected from a predefined set of emotion classes.
[0053] Alternatively, the control instruction further includes one or more of:
[0054] voice control instruction, configured to make selection among preset voice feedback templates and adjust voice parameters according to the emotion class;
[0055] atmosphere control instruction, configured to change one or more of brightness of a light source, color of a light source, flicker frequency of the light source, background music, based on the emotion class of the user.
[0056] Alternatively, the method for interactive controlling on adult product based on a multimodal model further includes:
[0057] performing multi-factor authentication and authorization management on user's identity to ensure that only authorized users control current adult product;
[0058] encrypting user data in a way of end-to-end to prevent unauthorized access and data leaks;
[0059] monitoring an execution process of the control instructions to block abnormal instructions when being detected, so as to ensure safe use.
[0060] Alternatively, the method for interactive controlling on adult product based on a multimodal model further includes:
[0061] downloading and installing a latest multimodal model parameters or control logic through a network connection or from a cloud platform, wherein the update process uses digital signatures and verification mechanisms to ensure authenticity and integrity of content being updated.
[0062] In a third aspect, the present disclosure further provides a device for interactive controlling on adult products based on a multimodal model, including:
[0063] a multimodal perception module configured to acquire perception data of a user interacting with the adult product, wherein the perception data of the user includes one or more modal data of voice data and tactile data;
[0064] a multimodal emotion recognition module configured to conduct feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto;
[0065] a multimodal fusion module configured to conduct feature-fusing on the extracted modal feature vectors to form cross-modal feature representation so as to determine emotion class of the user from a predefined set of emotion classes;
[0066] an instruction generation module configured to conduct search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product;
[0067] a control signal generation module configured to convert the control instruction into drive signals for the adult product, so that the adult product performs actions corresponding to the drive signals.
[0068] Alternatively, the instruction generation module conducting search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product includes:
[0069] conducting search in a data structure of action parameter vectors to generate a multi-component action parameter vector corresponding thereto, A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the specific emotion class of the user; the control instructions are configured to adjust one or more of the motion mode, intensity, or frequency of the adult product,
[0070] wherein, each component a1, a2, . . . , an represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions include one or more of vibration frequency, vibration intensity, contraction frequency, contraction intensity, swing amplitude, swing speed, temperature change amplitude, and pressure intensity of tactile feedback, a range of parameters of the physical control dimension corresponds to a preset set of movement patterns.
[0071] Alternatively, the control instruction further includes one or more of the following:
[0072] voice control instruction, configured to make selection among preset voice feedback templates and adjust voice parameters according to the emotion class;
[0073] atmosphere control instruction, configured to change one or more of brightness of a light source, color of a light source, flicker frequency of the light source, background music, based on the emotion class of the user
[0074] Alternatively, the device further includes:
[0075] an encryption module configured to perform multi-factor authentication and authorization management on user identity to ensure that only authorized users can control the current adult product; encrypt user data in a way of end-to-end to prevent unauthorized access and data leaks; monitor an execution process of the control instructions to block abnormal instructions to ensure safe use.
[0076] Alternatively, the device further includes:
[0077] a remote update module configured to download and install a latest multimodal model parameters or control logic through a network connection or from a cloud platform, where the update process uses digital signatures and verification mechanisms to ensure the authenticity and integrity of the updated content.
[0078] In a fourth aspect, the present disclosure further provides an adult product based on a multimodal model, including: a housing, a primary control component, a driving component, an execution component and an information acquisition component provided in or on the housing;
[0079] the information acquisition component includes at least one microphone and at least one sensor, the sensor includes at least one pressure sensor; the information acquisition is configured to acquire perception data of a user interacting with the adult product, wherein the perception data of the user includes one or more modal data of voice data and tactile data;
[0080] the primary control component is configured to process the perception data of the user by running a multimodal neural network model, the processing the perception data of the user includes: extracting voice feature vector from the voice data, and extracting tactile feature vector from a tactile data; applying an attention mechanism to the voice feature vector and the tactile feature vector to obtain modal weights by calculation and generate a weighted cross-modal feature representation that quantifies an emotional state of the user; mapping the weighted cross-modal feature representation to a specific emotion class of the user selected from a predefined set of emotion classes; conducting search in a data structure of action parameter vectors to generate a multi-component action parameter vector, A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the specific emotion class of the user; each component of the action parameter vector under searching represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions including at least vibration frequency and contraction intensity;
[0081] the driving component is configured to convert the numerical control values of the action parameter vector under searching into drive signals;
[0082] the execution component is configured to performs physical actions defined by the numerical control values.
[0083] Alternatively, the sensors include: a camera, a temperature sensor, a light sensor, and a sound sensor.
[0084] Alternatively, the driving component includes one or more of a motor, a servo motor, a pneumatic element, and a hydraulic element.
[0085] Alternatively, the execution component includes one or more of a microphone, a switch, a heater, a linear actuator, and a rotary actuator.
[0086] In a fifth aspect, the present disclosure provides a device for interactive controlling on adult products, including: at least one microphone and at least one sensor, one or more physical actuators; a processor; and a memory storing instructions thereon, when the instructions are executed by the processor, the device implements the steps of the methods as described above.
[0087] In a sixth aspect, the present disclosure provides an electronic apparatus including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the methods as described above.
[0088] In a seventh aspect, the present disclosure provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the method for interactive controlling on adult product based on a multimodal model.
[0089] Compared to the prior art, the embodiments of the present disclosure have at least the following advantageous effects.
[0090] The method for interactive controlling on adult product based on a multimodal model according to the embodiments of the present disclosure may achieve intelligent interaction by perceiving various modal data of the user, such as voice, text, images, and body movements with the multimodal model, to recognize the user's emotions accurately, and generating control instructions corresponding to the user's emotions, so that the adult product may be controlled to perform corresponding actions to implement intelligent interaction between the adult product and the user.
[0091] The method for interactive controlling on adult product based on a multimodal model according to the embodiments of the present disclosure addresses problems due to single interaction mode and unintelligent response in the prior art by using multimodal model technology so that higher interaction accuracy and adaptability may be achieved.BRIEF DESCRIPTION OF THE DRAWINGS
[0092] According to the description of non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more obvious:
[0093] FIG. 1 is a flowchart showing the method for interactive controlling on adult product based on a multimodal model according to an exemplary embodiment of the present disclosure;
[0094] FIG. 2 is a block diagram showing function units of the device for interactive controlling on adult product based on a multimodal model according to an exemplary embodiment of the present disclosure;
[0095] FIG. 3 is a block diagram showing the electronic apparatus according to an exemplary embodiment of the present disclosure.DETAILED DESCRIPTION
[0096] Detailed description would be made on the present disclosure in conjunction with specific embodiments. The following embodiments will facilitate understanding of skilled in the art on the present disclosure but place no limitation on the present disclosure in any way. It should be noted that various modifications and improvements can be made by skilled in the art without departing from the spirit of the present disclosure, which all fall within the scope of as claimed in the present disclosure.
[0097] It should be noted that the user information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, analyzed data, stored data, displayed data, etc.) involved in the present disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with relevant regulations.
[0098] The adult products described in the present disclosure may include various tangible adult products capable of interacting with users, including, but not limited to, massagers, vibrators, or other adult products with similar functions for female users; simulated vaginas, silicone dolls, or other adult products with similar functions for male users; and other adult products that can stimulate or influence users' sexual experiences and emotions. However, currently, an amplitude or frequency with which the adult product works may be adjusted manually or remotely during using the product, but cannot be adjusted according to sexual activities by users and thus the user may feel less interactions so that the user experience in using such products may be degraded and the user's intent expressed through language (text), voice (emotion), or body movements (intensity, frequency) cannot be understood properly, and the overall experience is degraded. Furthermore, there is no comprehensive processing on multimodal information (such as voice, text, images, body movements, and external environmental information) or methods for accurately recognizing user emotions using deep learning models (such as multimodal models) and generating corresponding control instructions in the prior art. Additionally, there are significant deficiencies in personalized user experience, real-time feedback, and safety assurance, and thus it is difficult to meet the higher demands for intelligent and emotional interaction. To address the above problems, the embodiments of the present disclosure provide a method for interactive controlling on adult product based on a multimodal model, which offers more intelligent interactions and provides users with a more immersive interactive experience.
[0099] In the following embodiments, in order to implement intelligent control for adult products (such as massagers, vibrators, or other adult products with similar functions for female users; simulated vaginas, silicone dolls, or other adult products with similar functions for male users; and other adult products that can stimulate or influence users' sexual experiences and emotions), information acquisition components such as sensors may be provided on the adult product itself or at the user. Additionally, controllable components such as motors, motion components, or driving components in specific parts of the adult product may be included. For example, sensors may be placed at each joint of silicone dolls to sense body movements. Driving components such as motors, servo motors, pneumatic elements, or hydraulic elements may be further included to control the movement of specific joints or components according to control instructions, so as to enable movements of parts of the silicone doll. Alternatively, the adult product may be connected to an app, which users may input voice, text, or image information thereby and may be provided on smart devices such as smartphones or smartwatches.
[0100] To achieve the above objectives, one embodiment of the present disclosure provides a method for dynamically controlling physical actuator of adult product, including:
[0101] a) acquiring perception data of a user interacting with the adult product by at least one microphone and at least one sensor, the perception data of the user includes one or more modal data of voice data and tactile data; the sensor includes at least one pressure sensor;
[0102] b) processing the perception data of the user by using a processor to running a multimodal neural network model, and particularly with the steps of 1) conducting feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto; 2) applying an attention mechanism to each modal feature vector to obtain modal weights by calculation and generate a weighted cross-modal feature representation that quantifies an emotional state of the user; 3) mapping the weighted cross-modal feature representation to a specific emotion class of the user selected from a predefined set of emotion classes;
[0103] c) conducting search in a data structure of action parameter vectors to generate a multi-component action parameter vector, A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the specific emotion class of the user; each component of the action parameter vector under searching represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions including at least vibration frequency and contraction intensity; and
[0104] d) converting the numerical control values of the action parameter vector under searching into drive signals and transmitting the drive signals to the physical actuator, so that the adult product performs physical actions defined by the numerical control values.
[0105] In the embodiments of the present disclosure, the control instruction further includes one or more of the following:
[0106] voice control instruction, configured to make selection among preset voice feedback templates and adjust voice parameters according to the emotion class;
[0107] atmosphere control instruction, configured to change one or more of brightness of a light source, color of a light source, flicker frequency of the light source, background music, based on the emotion class of the user.
[0108] In the embodiments of the present disclosure, the method for dynamically controlling physical actuator of adult product further includes:
[0109] acquiring real-time feedback data from the user, wherein the feedback data includes behavior response; and
[0110] dynamically adjusting predefined action parameter vectors stored in the data structure based on the real-time feedback data.
[0111] In the embodiments of the present disclosure, the method for dynamically controlling physical actuator of adult product further includes:
[0112] determining a priority of each control instruction based on a preset weight of the user's emotion class;
[0113] adjusting an execution sequence of a plurality of control instructions to ensure collaboration of the instructions and prevent conflicts therebetween;
[0114] acquiring real-time feedback data from the user including behavioral responses, and adjusting control strategies dynamically to optimize the user's experience.
[0115] In the embodiments of the present disclosure, the conducting feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto includes:
[0116] converting voice data using a pre-trained language model into voice text and extracting semantic representation vectors, when the perception data includes voice data;
[0117] extracting fundamental audio features from the voice data using Mel-frequency cepstral coefficients; extracting acoustic feature vectors from the fundamental audio features using spectrum analysis, the acoustic feature vectors include fundamental frequency, energy, and harmonic-to-noise ratio;
[0118] processing tactile data by using a multilayer perceptron network to obtain tactile features vectors corresponding to tactile data, when the perception data includes tactile data, the tactile feature vectors include one or more of touch pattern, touch force, touch rhythm, touch trajectory, touch area, and action frequency;
[0119] extracting facial expression features based on an improved FaceNet, when the perception data includes visual data, determining gazing duration and gazing frequency using a pupil tracking algorithm to obtain eye contact feature vectors; extracting body posture feature vectors using a human pose estimation model.
[0120] In the embodiments of the present disclosure, after determining emotion class of the user from a predefined set of emotion classes, the method further includes:
[0121] evaluating confidence of the emotion class by using a neural network.
[0122] In the embodiments of the present disclosure, the method for dynamically controlling physical actuator of adult product further includes:
[0123] forming temporal sequence features from the modal features corresponding to each type of modal data within a preset period and determining dynamic changes in emotion recognitions by processing the temporal sequence features using a long short-term memory network, to determine an emotional trend;
[0124] determining a moment of change in emotion, by focusing on transition points in the temporal sequence features with attention weights.
[0125] In the embodiments of the present disclosure, the method for dynamically controlling physical actuator of adult product further includes:
[0126] collecting user's interaction data and feedback information;
[0127] updating parameters of the multimodal model dynamically using deep learning algorithms and reinforcement learning algorithms;
[0128] adjusting control logic to meet user's personalized needs adaptively and optimizing determining on user intent and control strategies.
[0129] In the embodiments of the present disclosure, the method for dynamically controlling physical actuator of adult product further includes:
[0130] performing multi-factor authentication and authorization management on user's identity to ensure that only authorized users control current adult product;
[0131] encrypting user data in a way of end-to-end to prevent unauthorized access and data leaks;
[0132] monitoring an execution process of the control instructions to block abnormal instructions when being detected, so as to ensure safe use.
[0133] In the embodiments of the present disclosure, the method for dynamically controlling physical actuator of adult product further includes:
[0134] downloading and installing a latest multimodal model parameters or control logic through a network connection or from a cloud platform, wherein the update process uses digital signatures and verification mechanisms to ensure authenticity and integrity of content being updated.
[0135] Embodiments of the present disclosure provide a method for interactive controlling on adult product based on a multimodal model, starting with step S100.
[0136] S100. acquiring perception data of a user interacting with the adult product, wherein the perception data of the user includes one or more modal data of voice data and tactile data.
[0137] Specifically, in this step, information acquisition components such as sensors or sensing sheets may be provided on the adult product itself or at the user. Various technologies and devices may be used to acquire perception data such as voice, text, images, body movements, and external environmental information.
[0138] For example, voice data may be acquired via a microphone, and the acquired voice data may be further processed for recognition to be converted into text. Text information may be input via input devices. Visual data may include images or videos stored by the user or images or videos captured by a camera. Specifically, body movements may be obtained using high-speed cameras during tracking movements of human, or user movements may be obtained by using devices such as inertial measurement units, including, but not limited to, wearable devices (such as smartwatches). External environmental information may be acquired via sensors related thereto based on type of required information. For example, light sensors may acquire ambient brightness, temperature sensors may acquire ambient temperature, and sound sensors may acquire environmental sounds (such as music).
[0139] In the present embodiment, one or more types of perception data may be acquired as described above, with each type of data corresponding to a modal.
[0140] S200. conducting feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto.
[0141] In this step, feature extraction is conducted on each type of perception data to generate modal feature vectors corresponding thereto.
[0142] Specifically, feature extraction is conducted respectively for each type of perception data to generate single-modal feature representations. For example, feature extraction may be conducted on voice data to generate single-modal feature vectors corresponding to voice. Similarly, feature extraction may be conducted on tactile data to generate single-modal feature vectors corresponding to tactile data, and feature extraction may be conducted on visual data to generate single-modal feature vector corresponding to visual data.
[0143] S300. conducting feature-fusing on the extracted modal feature vectors to form cross-modal feature representation so as to determine emotion class of the user from a predefined set of emotion classes.
[0144] In this step, a multimodal model is used, and can use a prior model as the fundamental architecture. The multimodal model is pre-trained to recognize the user's current emotion class according to modal features.
[0145] After feature extraction is conducted on each type of perception data in step S200, a plurality of modal features, i.e., multimodal features, are obtained. Cross-attention processing is performed on the multimodal features to generate a unified feature representation.
[0146] In some embodiments, collected perception data may be used as training data in pre-training for the multimodal model, so that a multimodal model capable of recognizing user emotion class may be obtained. The feature representations obtained above are input into the trained multimodal model to recognize the user's emotion class.
[0147] In the present step, the user's emotion class may be recognized and determined based on one or more modal features in the perception data.
[0148] In embodiments, the user's emotion class may be set according to the practical situation of the adult product, and the user's emotion class may be recognized based on one or more types of perception data. The adult product mentioned in the present embodiment may be any adult product capable in providing interaction, including various shapes of adult dolls or other intelligent adult products.
[0149] S400. conducting search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product.
[0150] Control instructions may be generated according to o the user's emotion class, after the user's emotion class is determined in S300. Such control instructions may be configured to control the actionable or controllable parts of the adult product, to enable personalized intelligent interaction based on the user's needs and improve the user's experience. For example, the control instructions may include controlling the operating mode of the adult product, adjusting its working frequency, increasing or decreasing its motion amplitude, controlling expression of the adult product, providing feedback for voice chat, controlling movements of joints of the adult product, and other controllable parts that simulate human sexual behaviors.
[0151] S500. converting the control instruction into drive signals for the adult product, so that the adult product performs actions corresponding to the drive signals.
[0152] The control instructions correspond to specific user emotion class. The control instructions may be converted into drive signals suitable for the adult product when received in S400. These drive signals are transmitted to the control system or actuators of the adult product via physical interfaces (wired or wireless modules) to ensure that the drive signals effectively control or drive the corresponding parts of the adult product.
[0153] In the above embodiments of the present disclosure, with steps S100 to S500, feature representations may be obtained according to the information obtained from the user terminal and user's intent may be recognized by using the multimodal model and control instructions may be generated correspondingly according to the recognized user's intent. In the embodiments of the present disclosure, the multimodal model is applied in controlling on the adult products, i.e., human language, text, voice, emotions, body movements, and other data are perceived by using the multimodal model, so that user's intent may be recognized accurately, and the adult product may be controlled to respond accordingly based on the user's intent to form real-time intelligent interaction between the adult product and the user during use and improve the user's experience.
[0154] In using, although interaction requirements may be partially met by using single-modal features corresponding to a single modal, there may be instability caused by noise and other factors. To better recognize the user's emotion class, in step S200, user information may include two or more types of perception data.
[0155] In one embodiment of the method for interactive controlling on adult product based on a multimodal model according to the present disclosure, step S200 may include:
[0156] converting voice data using a pre-trained language model into voice text and extracting semantic representation vectors, when the perception data includes voice data; conducting recognition of emotional words and relationships using an emotional lexicon and semantic dependency analysis, the pre-trained language model is subjected to adaptive training in an adult field with respect to vocabulary and expressions in the adult field;
[0157] extracting fundamental audio features from the voice data using Mel-frequency cepstral coefficients; extracting acoustic feature vectors from the fundamental audio features using spectrum analysis, the acoustic feature vectors include fundamental frequency, energy, and harmonic-to-noise ratio; extracting high-dimensional voice representations using a self-supervised pre-trained speech encoder, to determine emotions and tones using an acoustic analyzer;
[0158] processing tactile data by using a multilayer perceptron network to extract tactile feature vectors when the perception data includes tactile data, the tactile feature vectors include one or more of touch pattern, touch force, touch rhythm, touch trajectory, touch area, and action frequency; and mapping the tactile features to a feature space related to emotions by nonlinear transformation.
[0159] extracting facial expression features based on an improved FaceNet, when the perception data includes visual data, determining gazing duration and gazing frequency using a pupil tracking algorithm to obtain eye contact feature vectors; extracting body posture features using a human pose estimation model; determining emotional states based on the facial expression features, eye contact feature vectors, and body posture features.
[0160] In this step, the perception data may be subjected to feature vectorization processing to generate feature vectors for emotional and sentiment processing and analysis.
[0161] In one embodiment of the method for interactive controlling on adult product based on a multimodal model according to the present disclosure, step S300 may include:
[0162] the step of conducting feature-fusing on the modal features to form cross-modal features to determine a current user's emotion class includes:
[0163] projecting each modal feature vector into a common dimensional feature space by linear mapping, to generate aligned modal features;
[0164] determining attention scores between the aligned modal features by calculation, to determine modal weights for generating a weighted cross-modal feature representation, and generate a weighted cross-modal feature representation that quantifies the emotional state of the user;
[0165] mapping the weighted cross-modal feature representation to a specific emotion class of the user selected from a predefined set of emotion classes.
[0166] In the embodiments of the present disclosure, first of all, one or more extracted modal features may be mapped through linear mapping and projecting aligned modal features into a feature space in a preset dimension;
[0167] secondly, correlation between a plurality of modal features may be determined by using an attention mechanism and conducting feature-fusing on the aligned modal features based on the correlation to form cross-modal feature representations;
[0168] at last, a distribution of possibilities for mapping the cross-modal features to various emotion classes may be determined by using a multilayer feedforward neural network and the current user's emotion class may be determined according to the probabilities.
[0169] The attention mechanism is used to determine the correlation between multiple modal features. Based on this correlation, the aligned modal features are fused to form cross-modal features.
[0170] A multilayer feedforward neural network is used to determine the probability distribution of the cross-modal features mapped to different emotion class, and the user's current emotion class is determined based on these probabilities.
[0171] The attention mechanism allows different parts of the features to be assigned with different weights automatically during processing, so as to better focus on and utilize important information. In this step, modal features are aligned by using the attention mechanism to learn how to focus on the correlation and importance between different modal features, and effectively capture the interrelationships between different modalities, so that overall performance and robustness may be improved. The attention mechanism may be implemented using neural networks to generate a series of attention vectors. These attention vectors represent key features in the input and are used to generate alignment results, so that modal features from different data sources remain consistent in terms of semantics and spatial dimensions.
[0172] In this step, the modal features are aligned for recognition of the user's emotion class, which may include mapping different modal features into same feature space, so that they may be understood and interacted with each other in this space.
[0173] In embodiments, the attention mechanism may be implemented using a Transformer model. The Transformer model may map features from different modalities into a shared feature space, so that features from different modalities may be compared and interacted with each other in the same space. This interaction may be achieved by using a cross-modal attention layer. Of course, in other embodiments, other attention mechanism networks may be used and there is no limitation thereon.
[0174] In some embodiments, to fuse the aligned features into cross-modal features, the step of fusing may further include: cascaded concatenation fusion, inter-modal relationship modeling, and deep feature extraction as described in detail below.
[0175] Cascaded concatenation fusion: directly concatenating all aligned modal features into a high-dimensional feature vector, to form initially fused cross-modal features, wherein the high-dimensional feature vector contains complete information from all modalities.
[0176] Inter-modal relationship modeling (attention mechanism): calculating attention mechanism the weight of each modal on the concatenated high-dimensional feature vector dynamically, and weighting concatenated features according to the modal weights to obtain weighted fused cross-modal features.
[0177] Deep feature extraction: inputting the weighted fused features into a deep learning network to further extract nonlinear relationships between modalities, and generating final cross-modal features.
[0178] In this embodiment, cascaded concatenation fusion, inter-modal relationship modeling, and deep feature extraction are used to fuse the aligned features into cross-modal features. With the combination of cascaded concatenation fusion and deep feature extraction, the maximum amount of information from different modalities is retained, and complex inter-modal interactions are effectively captured.
[0179] In the above embodiments of the present disclosure, the fusion of multimodal features may improve stability and generalization performance under different conditions, construct richer and more comprehensive semantic representations, and facilitate the understanding and processing of complex tasks, so that the intelligence of interaction between the adult product and the user may be further improved.
[0180] To further improve the classification of user emotion class, the user emotion class may further include emotional state quantification features. These features may be used to precisely control or adjust the control instructions of adult products during use by analyzing changes in emotional state quantification features. The emotional state quantification features refer to a quantitative expression of the user's emotion class, which can accurately reflect the user's emotion changes during using adult products. Emotional state quantification features may be achieved by adding a classifier during training of the neural network model to enable quantification and classification of emotion class. For instance, when the user's emotion class is recognized as “like,” the “like” emotion class may be quantified and represented using a feature vector, such as a scale from 0 to 10. A value of 0 indicates no emotion detected, 1-2 may represent a mild state, 3-5 may represent a normal state, 5-7 may represent a strong state, and 8-10 may represent an extreme state.
[0181] For example, when the user's current emotion class is determined as “like” and the emotional state quantification feature is 1-2, it can be considered as mild liking. In this case, the vibration intensity, frequency, and other parameters of the adult product may be controlled to be low. If the emotional state quantification feature is 5-7, it indicates strong liking, and the vibration intensity, frequency, motion amplitude, and other parameters of the adult product may be increased. Similarly, different settings may be applied to different modes and parameters of the adult product based on the emotional state quantification features correspondingly. For instance, an adult product may provide different feedback based on varying quantified features of the “like” emotional state within the 0-10 scale of emotional state quantification features. For example, when the emotional state quantification feature is 0, it may mean a neutral expression with no action. When the feature is 1-2, it means mild liking, and the product may show a slight smile. When the feature is 3-5, it means a normal state of “like,” and the product may show a noticeable smile (with a larger curvature of the mouth) with slight movements (such as nodding and / or body movements). When the feature is 5-7, it means a strong state of “like,” and the product may show a laughing expression with larger body movements. Correspondingly, when the feature is 8-10, the product may show an exuberant laugh along with more expressions and larger body movements, to implement precise interactions in terms of expressions and gestures similar to those of a real person.
[0182] It should be noted that all control or parameter adjustments of the adult product are conducted within a safety range of the product. The above examples are for illustrative purposes only. In other embodiments, other classification methods and expressions may be used according to the specific usage scenarios and conditions of the adult product.
[0183] In some preferred embodiments, after the user's emotion class is recognized in step S300, the corresponding step S400 of conducting search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product includes:
[0184] conducting search in a data structure of action parameter vectors to generate a multi-component action parameter vector corresponding thereto, A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the specific emotion class of the user; the control instructions are configured to adjust one or more of the motion mode, intensity, or frequency of the adult product,
[0185] each component a1, a2, . . . , an represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions include one or more of vibration frequency, vibration intensity, contraction frequency, contraction intensity, swing amplitude, swing speed, temperature change amplitude, and pressure intensity of tactile feedback, a range of parameters of the physical control dimension corresponds to a preset set of movement patterns.
[0186] In the embodiments of the present disclosure, a control instruction may be generated according to the emotion class, wherein the control instruction is configured to change one or more of a movement mode, an intensity, or a frequency in which the adult product is operating. The control instruction is represented as an action parameter vector A=[a1, a2, . . . , an].
[0187] More particularly, each of the components a1, a2, . . . , an represents an independent physical control dimension for the adult product. The physical control dimension may include one or more of the following: vibration frequency, vibration intensity, contraction frequency, contraction intensity, swing amplitude, swing speed, temperature change amplitude, and pressure intensity of tactile feedback. A range of parameters of the physical control dimension corresponds to a preset set of movement patterns.
[0188] When the emotion class is defined with the above-mentioned emotional state quantification features, the corresponding process in step S400 for generating control instructions includes:
[0189] analyzing the emotional state quantification features to generate control parameters for one or more of motion mode, intensity, or frequency of the adult product and generating control instructions.
[0190] In the above embodiments of the present disclosure, control instructions and parameters are generated correspondingly based on recognized user's current emotion class, emotional state quantification features, and external environmental information, so that the user's experience may be optimized, and interaction intelligence and accuracy may be significantly improved.
[0191] The above preferred embodiments further define the user's emotion class and emotional state level by using emotional state quantification features. During the interaction between the adult product and the user, the emotional state quantification features allow precise control or adjustment of the state and / or parameters during use of the adult products, to provide better emotional feedback for the user, including voice interaction feedback, adjustments to the working modes and parameters of the adult product, and humanoid operational feedback from the adult product. This approach is more intelligent and accurate than the solution of merely interacting based on a single emotional recognition and meets user needs more.
[0192] In the above embodiments of the present disclosure, after the user's emotion class (including emotional state quantification features) is recognized accurately by using the multimodal model, control instructions are generated correspondingly based on the recognized emotion class. In some embodiments, the control instructions are represented as an action parameter vector A=[a1, a2, . . . , an].
[0193] Each of the components a1, a2, . . . , an represents an independent physical control dimension for the adult product. The physical control dimension may include one or more of the following: vibration frequency, vibration intensity, contraction frequency, contraction intensity, swing amplitude, swing speed, temperature change amplitude, and pressure intensity of tactile feedback. A range of parameters of the physical control dimension corresponds to a preset set of movement patterns.
[0194] The control instruction further includes voice control instruction, atmosphere control instruction, or the like.
[0195] More particularly, the control instruction may be configured to adjust one or more of the motion mode, intensity, or frequency of the adult product. The voice control instruction may make selection among preset voice feedback templates and adjust voice parameters (such as voice tune or content) according to the emotion class. A range of parameters of the motion mode corresponds to a preset set of movement patterns and may be determined according to type or structure of the adult product. Intensity or frequency may refer to the intensity or frequency of vibrations, movements, etc., of the adult product, as well as the intensity or frequency of the product's action feedback or voice response. For example, if the feedback expression is “like,” it may be a mild like or a strong like. The adult product may perform different interactions with the user in terms of motion and voice under different emotional features. The entire process simulates reactions of real human as closely as possible in tone or actions (such as movements of human-like joints) to create a realistic interactive scenario.
[0196] Furthermore, atmosphere control instruction for adult products may be further included to optimize the atmosphere where the user uses the adult product. The atmosphere control instruction is configured to change one or more of brightness of a light source, color of a light source, flicker frequency of the light source, background music, based on the emotion class of the user. The atmosphere control instruction in the embodiments of the present disclosure may be configured to change one or more of brightness, color, flicker frequency of lighting, background music to match the user's emotions.
[0197] In the embodiments of the present disclosure, real-time feedback data may be acquired from the user and the feedback data may include behavioral responses, and control strategies may be dynamically adjusted according the real-time feedback data to optimize the user's experience.
[0198] In the above embodiments of the present disclosure, intelligent interaction may be achieved accurately by recognizing the user's emotions and generating corresponding control instructions. Furthermore, to provide more options and better user's experience, in some embodiments, the method for interactive controlling on adult product based on a multimodal model according to the present disclosure further includes scheduling the control instructions, which may include determining a priority of each control instruction based on a preset weight of the user's emotion class; adjusting an execution sequence of a plurality of control instructions to ensure collaboration of the instructions and prevent conflicts therebetween. Furthermore, the method according to the present disclosure may further include: acquiring real-time feedback data from the user, including behavioral responses, and adjusting control strategies dynamically in combined with environmental information and emotional state quantification features to optimize the user experience.
[0199] In the above embodiments, description is made on recognizing user emotion recognition and generating control instruction. For better understanding, in some embodiments, the process of generating control instructions based on the recognized user emotion class and emotional state quantification features may include:
[0200] M1. determining a target control type, such as motion, voice, or atmosphere control.
[0201] M2. determining preset control parameters based on the user's emotional state and / or external environmental information (such as light brightness, background music, temperature, etc.).
[0202] M3. adjusting control parameters, such as motion intensity, frequency, voice volume and tone, light brightness and color, etc., dynamically, to generate an action parameter vector, so that the control instructions remain consistent with the user's emotional state and / or environmental state.
[0203] M4. generating a final comprehensive control instructions and further optimizing the instruction parameters by using a real-time feedback mechanism (such as user behavioral responses).
[0204] It should be noted that in some embodiments, the process of generating control instructions may further include:
[0205] (1) determining the user's emotion level (such as relaxed, excited, or tense) based on the emotional state quantification features and quantifying the corresponding value range of the action parameter vector A=[a1, a2, . . . , an];
[0206] (2) analyzing external environmental information (such as light brightness, background music, temperature, etc.) and determining a preset optimization strategy based on the emotional state quantification features;
[0207] (3) adjusting the parameters of the action parameter vector A=[a1, a2, . . . , an](such as vibration intensity range, light color and brightness, volume, etc.) dynamically so that the control instructions remain consistent with the user's emotional and environmental states.
[0208] (4) generating a final comprehensive control instruction and further optimizing the instruction parameters via a real-time feedback mechanism (such as user behavioral responses).
[0209] According to the above embodiments of the present disclosure, to continuously improve the control on adult products during use, in some embodiments, the method for interactive controlling on adult product based on a multimodal model may further optimize user emotion recognition and control strategies. Specifically, the method for interactive controlling on adult product based on a multimodal model may include the following optimizations for machine learning: collecting user interaction data and feedback information; updating parameters of the multimodal model dynamically using deep learning algorithms and reinforcement learning algorithms; adjusting control logic to meet user's personalized needs adaptively and optimizing determining on user intent and control strategies to improve the intelligence of the system.
[0210] To ensure security of user data and protect user privacy, according to the above embodiments of the present disclosure, the method for interactive controlling on adult product based on a multimodal model may further include multi-factor authentication and authorization management for user identity, so that only authorized users may control the adult product. Additionally, user data may be processed with end-to-end encryption to prevent unauthorized access and data breaches. Furthermore, the execution process of control instructions may be monitored in real-time to detect and block abnormal instructions, ensuring safe usage.
[0211] According to the above embodiments of the present disclosure, the method for interactive controlling on adult product based on a multimodal model may further include downloading and installing a latest multimodal model parameters or control logic through a network connection or from a cloud platform, where the update process uses digital signatures and verification mechanisms to ensure the authenticity and integrity of the updated content.
[0212] Referring to FIG. 2, to better recognize user's intent, in another embodiment, when user's information is of multiple types (two or more) such as voice, text, image, body movements, and external environmental information, the multimodal model may includes two or more of a speech understanding unit, a voice emotion recognition unit, a facial expression recognition unit, a motion recognition unit, and a context understanding unit. The multimodal model may further include a multimodal semantic fusion unit. The speech understanding unit, voice emotion recognition unit, facial expression recognition unit, motion recognition unit, and context understanding unit correspond to each of unimodal features, while the multimodal semantic fusion unit performs semantic fusion on multiple modal features to obtain the fused multimodal features.
[0213] The method for interactive controlling on adult product based on a multimodal model according to the embodiments of the present disclosure may implement personalized feedback. By dynamically adjusting control parameters based on the user's emotional state quantification features and external environmental information, personalized motion, voice, or atmosphere feedback may be provided according to the user's real-time needs, so that the immersion and realism of the user experience are significantly improved.
[0214] The method for interactive controlling on adult product based on a multimodal model according to the embodiments of the present disclosure provides security in system with security guarantees such as multi-factor authentication, data encryption, and abnormal instruction detection mechanisms, so that the security of user data and the device control process may be ensured to further improve user trust.
[0215] Specifically, the functional units implement the following functions:
[0216] The speech understanding unit is configured to analyze and understand the semantic features of the text in the user's information.
[0217] The voice emotion recognition unit is configured to analyze emotional features of a voice in the user's information.
[0218] The facial expression recognition unit is configured to analyze semantic features of facial expressions in an image of the user's information.
[0219] The motion recognition unit is configured to understand semantics of body movements in the user's information.
[0220] The context understanding unit is configured to construct a semantic understanding framework by combining historical interaction information and current dialogue content.
[0221] The multimodal semantic fusion unit is configured to integrate the semantic analyzing results of the functional units to generate unified semantic understanding, when two or more of the above functional units are included.
[0222] In the above embodiments, the language understanding unit may employ various natural language processing (NLP) techniques, such as pre-trained deep learning models, to parse and understand the semantic features of the text in user's information based on sentence structure, semantic analysis, etc.
[0223] In the above embodiments, the voice emotion recognition unit may extract voice-related features and speech-related features. Voice-related features may include fundamental frequency and energy. Fundamental frequency reflects a frequency of the pitch in a speech signal and is related to the speaker's pitch. Energy reflects an intensity of a speech signal and is related to the speaker's volume and emotional state. Speech-related features may include speech speed, intonation, and spectral features. Speech speed reflects a speed of speech, which is related to the emotional state. Intonation reflects changes in pitch in a speech signal and is closely related to the emotional state. Spectral features reflect distribution characteristics of a speech signal on the spectrum and are related to pronunciation habits and emotional state. With these features and further extracted by using machine learning or deep learning techniques to be used in training an emotion recognition model (neural network), analyzing may be conducted on the emotional features of the voice in the user's information.
[0224] In the above embodiments, the facial expression recognition unit generally first detects a face in an image and then extracts key feature information of the face, such as shape, size, and position of eyes, mouth, eyebrows, etc. The way for extracting features may include geometry-based extraction and deep learning-based extraction (such as CNN). Classification and recognition on expressions may be implemented based on detected key feature information. Common classifiers may include SVM (support vector machines), decision trees, neural networks, and deep learning models. Of course, before being used for classification, the model may be trained and tested using image datasets with known expression labels to train the classifier and verify its accuracy and robustness by using test sets. Test images may be directly input into pre-trained models to obtain expression classification results, so as to complete the analysis on the semantic features of facial expressions in user's information.
[0225] In the above embodiments, the motion recognition unit, similar to the above units, may obtain user images and achieve motion recognition using pre-trained motion recognition models or keypoint detection means. Motion recognition may be further conducted by using data or information obtained from sensing devices. For example, for users of adult products, wearable devices (such as wristbands) or other sensing devices (such as three-axis or six-axis sensors) may conduct detection to determine the speed, direction, and force magnitude of the user's body movements. Based on the results of motion recognition, the semantics of body movements in the user's information may be determined.
[0226] In the above embodiments, the context understanding unit may process historical interaction information and current dialogue content, which may include voice, text, or image information. With analyzing conducted on the semantic relationships between words (such as synonyms, antonyms, hypernyms, and hyponyms) and the logical relationships between sentences (such as causal, adversative, and other relationships), and incorporating contextual information, a semantic understanding framework may be constructed so that the current dialogue content may be determined more accurately to generate context semantics. Furthermore, the semantic understanding framework may be continuously optimized and adjusted by accumulating historical interaction information and learning different user habits and preferences, so as to provide a better understanding framework for subsequent user intent recognition.
[0227] It can be understood that the context understanding unit may integrate historical interaction information and current dialogue content, analyze semantic relationships and logical structures, and generate contextual semantic features that reflect the user's current intent and historical behavior. These semantic features have stronger contextual relevance and dynamic adaptability.
[0228] In the case that two or more of the above units are used, the multimodal semantic fusion unit integrates semantic analyzing results therefrom to generate unified semantic understanding. That is, the multimodal semantic fusion unit combines the semantic analyzing results from different sources or modes (such as voice, text, image, body movements, and external environmental information) to obtain a richer and more comprehensive unified semantic understanding. For example, the multimodal semantic fusion unit may use deep learning models to achieve feature-level fusion by using shared layers or joint training. Alternatively, each modal may independently make predictions, whose results may be combined at the decision stage. A multimodal neural network architecture may be constructed, with additional layers responsible for cross-modal feature fusion, and unified semantic understanding may be finally obtained.
[0229] According to the technical idea for the method for interactive controlling on adult product based on a multimodal model described above, another embodiment of the present disclosure may further provide a device for interactive controlling on adult product based on a multimodal model, including:
[0230] a multimodal perception module configured to acquire perception data of a user interacting with the adult product, and the perception data of the user includes one or more modal data of voice data and tactile data;
[0231] a multimodal emotion recognition module configured to conduct feature-extracting on each modal data in the perception data to obtain modal feature vectors corresponding thereto;
[0232] a multimodal fusion module configured to conduct conducting search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product;
[0233] an instruction generation module configured to conduct search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product;
[0234] a control signal generation module configured to convert the control instruction into drive signals for the adult product, so that the adult product performs actions corresponding to the drive signals.
[0235] In some embodiments of the present disclosure, the multimodal perception module may obtain various types of information through one or more interfaces. Some of the obtained information may be subjected to preprocessing (such as preliminary data formatting, noise reduction, etc.), before being transmitted to the multimodal emotion recognition module. Furthermore, perception data may be collected by corresponding sensors or devices (such as microphones, cameras, tactile force feedback devices, etc.) and may be obtained in a wired or wireless way.
[0236] In some embodiments, more particularly, the multimodal emotion recognition module extracts feature vectors (such as voice emotion features, text semantic features, facial expression features, etc.) and performs emotion recognition on each type of features.
[0237] In the device according to embodiments of the present disclosure, the instruction generation module is further configured to generate a control instruction according to the emotion class, wherein the control instruction is configured to change one or more of a movement mode, an intensity, or a frequency in which the adult product is operating. The control instruction is represented as an action parameter vector A=[a1, a2, . . . , an], where each of the components a1, a2, . . . , an represents an independent physical control dimension for the adult product. The physical control dimension may include one or more of the following: vibration frequency, vibration intensity, contraction frequency, contraction intensity, swing amplitude, swing speed, temperature change amplitude, and pressure intensity of tactile feedback. A range of parameters of the physical control dimension corresponds to a preset set of movement patterns.
[0238] In the device according to embodiments of the present disclosure, the control instruction further includes one or more of the following:
[0239] voice control instruction, configured to make selection among preset voice feedback templates and adjust voice parameters according to the emotion class.
[0240] atmosphere control instruction, configured to change one or more of brightness of a light source, color of a light source, flicker frequency of the light source, background music, based on the emotion class of the user.
[0241] The device according to embodiments of the present disclosure further includes:
[0242] an encryption module configured to perform multi-factor authentication and authorization management on user identity to ensure that only authorized users can control the current adult product; encrypt user data in a way of end-to-end to prevent unauthorized access and data leaks; monitor an execution process of the control instructions to block abnormal instructions to ensure safe use.
[0243] The device according to embodiments of the present disclosure further includes:
[0244] a remote update module configured to download and install a latest multimodal model parameters or control logic through a network connection or from a cloud platform, where the update process uses digital signatures and verification mechanisms to ensure the authenticity and integrity of the updated content.Embodiment 1: The user is male, and the adult product is a smart silicone doll
[0245] User behavior: The user says to the smart silicone doll, “You look beautiful today”, while touching the doll's cheek gently and slowly.
[0246] Emotion recognition: four input signals may be processed by the multimodal Transformer model: voice content (converted into text via speech recognition), voice acoustic features (such as fundamental frequency, energy, rhythm, etc.), touch pressure data (collected through a pressure sensor array), and touch speed / trajectory data (obtained in real-time by capacitive sensors). Each modal data is first processed by an independent feature extractor and then fused across modalities by using a self-attention mechanism to capture correlations between different signals. The model adopts a pretraining-finetuning paradigm to be pretrained on large-scale emotional interaction data and then finetuned for specific scenarios. Confidence scores may be dynamically adjusted based on historical interactions by using an adaptive threshold mechanism.
[0247] The model ultimately outputs an emotion class of “gentle / intimate,” with a confidence score of 85%, along with the analyzing results for emotional intensity and temporal trends.Response Made by the Doll:
[0248] The doll's facial expression changes to a smile, and its eyes perform a slow blinking motion.
[0249] The facial area lightly heats up, while simulating a blushing reaction.
[0250] The doll responds by using its built-in speaker: “Thank you, I think you look especially charming today” with a soft tone of voice.
[0251] The doll's body posture changes to gently lean toward the user.Additional Feedback (Optional):
[0252] The ambient lighting around the doll automatically dims to create a romantic atmosphere.
[0253] A notification is displayed on the mobile application indicating that “Intimacy Mode has been activated,” while playing soft music as background music.
[0254] The system records this interaction data for future optimization for response patterns.Embodiment 2: The User is Female, and the Adult Product is a Smart Massager
[0255] User behavior: During use, the user softly says, “Can you go faster,” or makes weak vocalizations such as moaning, while slightly accelerating the movement of the massager by her wrist.
[0256] Emotion recognition: The multimodal Transformer architecture employs a specially enhanced processing flow designed for low signal-to-noise ratio environments. The voice processing branch uses an acoustic model specifically trained for weak sounds and non-verbal vocalizations (such as moaning and panting) by self-supervised learning means to extract feature patterns from large-scale contextual audio. The model employs a specialized attention mechanism capable of capturing subtle sound variations below 20 dB in noisy environments, with spectral enhancement techniques to amplify critical frequency bands.
[0257] High-precision millimeter-wave sensor array capable of recognizing subtle acceleration changes as low as 0.2 mm / s2 may be used for the motion recognition branch, combined with pressure sensors to capture slight pressure changes in the range of 1-5 g. A temporal Transformer employs a sliding window technique to perform cumulative analysis on micro-movements over the past 30 seconds, to recognize potential rhythm change trends.
[0258] The two modalities are integrated by using a confidence-weighted fusion mechanism. When one modal signal is weak, the system automatically increases the weight of the other modal. The system ultimately accurately identifies the “excited / anticipating enhancement” state from extremely subtle moaning and movements, of hands, with a confidence level of 93%, and provides sensitivity adjustment options to allow the system to dynamically adjust recognition thresholds based on the usage scenario.Adjustments to Massager:
[0259] The vibration frequency gradually increases from the 65 Hz to 90 Hz.
[0260] The amplitude and intensity gradually increase from 60% to 75%.
[0261] The vibration mode transitions from steady mode to a wave-like fluctuation mode.Adjusting Process:
[0262] The frequency is increased by 15 Hz within the first 3 seconds.
[0263] The frequency and amplitude may be changed periodically in a wave-like curve during the next 5 seconds.
[0264] Fine-tuning is performed every 20 seconds, based on the user's reactions (voice, motion frequency) to optimize vibration parameters automatically.
[0265] Keeping monitoring the user's reactions, and when changes in the user's voice or movements are detected, the parameters are automatically adjusted to match emotional changes.Feedbacks Along with the Adjustments (Optional):
[0266] The massager's temperature gradually increases by 2° C. above body temperature.
[0267] A notification tone and voice message may be made by a smartphone connected with the massager: “Enhancement Mode activated”.
[0268] A smart lighting system connected with the massager changes brightness rhythmically in sync with the vibration frequency.
[0269] An App records preference data during the using of the massager for personalized recommendations in future.
[0270] The embodiments of the present disclosure further provide an adult product based on a multimodal model, including: a housing, a primary control component, a driving component, an execution component and an information acquisition component provided in or on the housing;
[0271] the information acquisition component includes at least one microphone and at least one sensor, the sensor includes at least one pressure sensor; the information acquisition is configured to acquire perception data of a user interacting with the adult product, wherein the perception data of the user includes one or more modal data of voice data and tactile data;
[0272] the primary control component is configured to process the perception data of the user by running a multimodal neural network model, the processing the perception data of the user includes: extracting voice feature vector from the voice data, and extracting tactile feature vector from a tactile data; applying an attention mechanism to the voice feature vector and the tactile feature vector to obtain modal weights by calculation and generate a weighted cross-modal feature representation that quantifies an emotional state of the user; mapping the weighted cross-modal feature representation to a specific emotion class of the user selected from a predefined set of emotion classes; conducting search in a data structure of action parameter vectors to generate a multi-component action parameter vector, A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the specific emotion class of the user; each component of the action parameter vector under searching represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions including at least vibration frequency and contraction intensity;
[0273] the driving component is configured to convert the numerical control values of the action parameter vector under searching into drive signals;
[0274] the execution component is configured to performs physical actions defined by the numerical control values.
[0275] The drive signals generated in the embodiments of the present disclosure match the model of the execution component used. When control instructions are converted into drive signals, the command set adapted to the model of the execution component is used, and the drive signals are generated according to the gear levels and operating modes supported by the model of the execution component. In embodiments, the gear levels and operating modes supported by different motors may vary. For example, motor a supports three levels of speed, while motor b supports five levels of speed. When drive signals for motor a and motor b are generated, the conversion is performed in a nearest-match manner to the level of speed supported by motor a and motor b, respectively.
[0276] Alternatively, the sensors include: a camera, a temperature sensor, a light sensor, and a sound sensor.
[0277] Alternatively, the driving component includes one or more of a motor, a servo motor, a pneumatic element, and a hydraulic element.
[0278] Alternatively, the execution component includes one or more of a microphone, a switch, a heater, a linear actuator, and a rotary actuator.Embodiment 3: The User is Male, and the Adult Product is a Smart Silicone Doll or a Device with a Similar Function
[0279] Pressure sensors are provided on chest and other locations of the smart silicone doll. Temperature sensors are provided on the genital area and other locations of the smart silicone doll. Other types of sensors, such as cameras, light sensors, or sound sensors, are provided in positions as needed on the smart silicone doll. A smart terminal may be used as a camera, light sensor, or sound sensor.
[0280] Rotary actuators and linear actuators are provided at each joint and movable part of the smart silicone doll. Each actuator is connected to a motor or driving component, and a clamping mechanism is provided in the genital area.
[0281] The main control component conducts analyzing on one or more types of collected perception data and determines control instructions according to the above method, so as to control one or more of eyes, mouth, head, shoulders, elbows, wrists, hips, knees, ankles, genitals, and buttocks to perform corresponding actions. These actions are combined with heating, voice, and other feedback to achieve intelligent interaction between the adult product and the user.
[0282] Specifically, when the eyes, mouth, head, shoulders, elbows, wrists, hips, knees, ankles, genitals, and buttocks perform actions, the amplitude or intensity of the actions may be set to several states. The number of states set for each part may be the same or different. For example, the eyes may be set to two states: open and closed. The head may be set to nine states: looking forward, turning 100 to the left, turning 20° to the left, turning 30° to the left, turning 40° to the left, turning 100 to the right, turning 200 to the right, turning 300 to the right, and turning 400 to the right. The state settings for each part may be as shown in Table 1:
[0283] TABLE 1partstate 1state 2state 3state 4state 5state 6state 7state 8state 9eyesOpenSquintClosedMouthFullyOpenOpenOpenClosedOpenLevel 1Level 2Level 3HeadLooking-Turning-Turn-Turn-Turn-Turn-Turn-Turn-Turn-ForwardLeftLeftLeftLeftRightRightRightRight1Level 2Level 3Level 4Level 1Level 2Level 3Level 4ShoulderDroppedUPUPUPUPUPUPUPFullyLevel 1Level 2Level 3Level 4Level 5Level 6Level 7UPElbowStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8WristStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8Hips (Open)OpenOpenOpenOpenOpenOpenOpenOpenOpenLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8Level 9Hips (Bent)StraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8KneesStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8AnklesStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8GenitalsA1A2A3A4A5A6A7(Frequency)times / mintimes / times / times / times / times / times / minminminminminminGenitalsLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8Level 9(Amplitude)ButtocksAmplitude Amplitude Amplitude Amplitude Amplitude Amplitude Amplitude Amplitude Amplitude (Amplitude)123456789ButtocksB1B2B3B4B5B6B7(Frequency)times / mintimes / times / times / times / times / times / minminminminminmin. . .
[0284] Movements of arms may be controlled by controlling states of the shoulders, elbows, and wrist, so as to hug the user and may be further controlled to control the strength of the hug.
[0285] Movements of legs may be controlled to bend by controlling states of the hips, knees, and ankles, and may be further controlled to control the degree of bending.
[0286] The genital may be controlled to be tight by controlling action thereof, and may be further controlled to control frequency and intensity of tightening.
[0287] The buttocks may be controlled to swing by controlling action thereof, and may be further controlled to control the amplitude and speed of swinging.Embodiment 4: The User is Female, and the Adult Product is a Male Doll, a Massager, or a Device with a Similar Function
[0288] Pressure sensors are provided on the genital area and other parts of the smart male doll. Temperature sensors are provided on the genital area and other parts of the smart male doll. Other types of sensors, such as cameras, light sensors, or sound sensors, are provided as needed on the smart male doll. A smart terminal may be alternatively used as a camera, light sensor, or sound sensor.
[0289] Rotary actuators and linear actuators are provided at each joint and movable part of the smart male doll. Each actuator is connected to a motor or driving component, and a vibration mechanism is provided in the genital area.
[0290] The main control component conducts analyzing on one or more types of collected perception data and determines control instructions according to the above method, so as to control one or more of eyes, mouth, head, shoulders, elbows, wrists, hips, knees, ankles, genitals, and buttocks to perform corresponding actions. These actions are combined with heating, voice, and other feedback to conduct intelligent interaction between the adult product and the user.
[0291] More particularly, when the eyes, mouth, head, shoulders, elbows, wrists, hips, knees, ankles, genitals, and buttocks perform actions, the amplitude or intensity of the actions may be provided with several states. The state for each part may be as shown in Table 2.
[0292] TABLE 2partstate 1state 2state 3state 4state 5state 6state 7state 8state 9eyesOpenSquintClosedMouthFullyOpenOpenOpenClosedOpenLevel 1Level 2Level 3HeadLooking-Turning-Turn-Turn-Turn-Turn-Turn-Turn-Turn-ForwardLeftLeftLeftLeftRightRightRightRight1Level 2Level 3Level 4Level 1Level 2Level 3Level 4ShoulderDroppedUPUPUPUPUPUPUPFullyLevel 1Level 2Level 3Level 4Level 5Level 6Level 7UPElbowStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8WristStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8Hips (Open)OpenOpenOpenOpenOpenOpenOpenOpenOpenLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8Level 9Hips (Bent)StraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8KneesStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8AnklesStraightBentBentBentBentBentBentBentBentLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8GenitalsA1A2A3A4A5A6A7(Frequency)times / mintimes / times / times / times / times / times / minminminminminminGenitalsLevel 1Level 2Level 3Level 4Level 5Level 6Level 7Level 8Level 9(Amplitude)ButtocksAmplitude Amplitude Amplitude Amplitude Amplitude Amplitude Amplitude Amplitude Amplitude (Amplitude)123456789ButtocksB1B2B3B4B5B6B7(Frequency)times / mintimes / times / times / times / times / times / minminminminminmin. . .
[0293] Movements of arms may be controlled by controlling states of the shoulders, elbows, and wrist, so as to hug the user and may be further controlled to control the strength of the hug.
[0294] Movements of legs may be controlled to bend by controlling states of the hips, knees, and ankles, and may be further controlled to control the degree of bending.
[0295] The genital may be controlled to be tight by controlling action thereof, and may be further controlled to control frequency and intensity of tightening.
[0296] The buttocks may be controlled to swing by controlling action thereof, and may be further controlled to control the amplitude and speed of swinging.
[0297] Specifically, the present embodiment may include the following.System Architecture
[0298] A system may be embedded into adult products, and may process a plurality of input data, such as speech, contact force, and vision, and conduct emotion recognition and feedback by using deep learning means.Multimodal Perception Module
[0299] The system collects the following input data:
[0300] 1. Voice data: voice content and acoustic features are collected through a microphone array.
[0301] 2. Tactile data: contact force, speed, and trajectory are obtained through a pressure sensor array.
[0302] 3. Visual data (optional): facial expressions and body movements are captured through a camera.Algorithm for Multimodal Emotion Recognition1. Feature Extraction LayerDedicated Feature Extraction Sub-Networks are Configured for Input Data of Each Modal:
[0303] Voice content feature extraction: A Chinese pre-trained language model (such as the BERT model) is used to process voice text obtained by conversion and extract semantic representations. With analyzing on emotional lexicon features and semantic dependency, the recognition on emotional words and relationships may be improved.
[0304] Field adaptation training is conducted for specific vocabulary and expressions in the adult filed.
[0305] Voice acoustic feature extraction: Mel-frequency cepstral coefficients (MFCC) are used to extract basic acoustic features. Spectral analysis is employed to extract fundamental frequency, energy, harmonic-to-noise ratio, and other acoustic features. High-dimensional voice representations are extracted by using a self-supervised pre-trained speech encoder (such as a variant of wav2vec 2.0). Spectral enhancement techniques are used in low signal-to-noise ratio environments to improve the capture of signals below 20 dB.
[0306] Tactile feature extraction: A multilayer perceptron network is used to process touch and pressure sensor data to extract touch patterns, force, and rhythm features. Nonlinear transformations map raw tactile signals into emotion-related feature spaces. Pressure time-series features are collected in a range of 1-500 g with a sampling rate of 100 Hz. Touch trajectory features are extracted using a spatiotemporal convolutional network. Touch area distribution is obtained through a pressure sensor array to generate a contact area heat map. Action frequency features are extracted by using FFT analysis to identify periodic characteristics.
[0307] Visual feature extraction (if applicable): Facial expression features are extracted based on an improved FaceNet to obtain features of 68 key points. Eye contact features are determined by using a pupil tracking algorithm to determine gaze duration and frequency. Body posture features are extracted using a lightweight human pose estimation model to identify key points.2. Multimodal Fusion Layera Multimodal Transformer Architecture is Used for Interaction and Fusion Between Modalities:
[0308] Feature mapping: Features of each modal are projected into same dimensional feature space by using linear mapping to create a unified representation basis.
[0309] Multi-head attention mechanism: A self-attention mechanism may be used for various modalities to focus on and learn from each other. For example, there may be a correlation between voice intonation and touch force, which can be captured by using the attention mechanism.
[0310] Modal interaction: Information exchange and complementation between modalities are adopted. When one modal signal is unclear, supplementary information may be obtained from other modalities. For example, when the voice signal is weak, the system pays more attention on changes in tactile signals.
[0311] Residual connections and layer normalization: Effective information transmission may be implemented in deep networks by using residual connections and layer normalization, so as to avoid gradient vanishing problems, and maintain numerical stability.3. Emotion Classification and Confidence EvaluationEmotion State Recognition is Performed Based on the Fused Representation:
[0312] Emotion classifier: A multilayer feedforward neural network is used to map the fused features to the probability distribution of different emotion class. A specific emotion class system is defined for adult product scenarios, including specific emotional states such as “gentle / intimate” and “excited / anticipating enhancement.”
[0313] Confidence estimator: An independent neural network branch may be used to evaluate reliability of emotion recognition results. When an input data has much noise or is ambiguous, a lower confidence score is output.
[0314] Emotion intensity estimator: Evaluation may be conducted on intensity of the emotion to output a value between 0 and 1 to represent the strength of the emotion.
[0315] Intensity information is used to adjust the amplitude of device in response.4. Adaptive Threshold MechanismThe Recognition Threshold is Dynamically Adjusted for Different Users and Scenarios:
[0316] History data accumulation: Confidence scores from the most recent 10 interactions are recorded to establish an adjustment basis. The system pays attention on the user's interaction patterns in history to be adapted to individual usage habits.
[0317] Threshold adaptive adjustment: Based on history confidence and user's potential feedback, decision thresholds are dynamically adjusted. For example, when the system repeatedly has high confidence, the threshold may be lowered appropriately to respond more actively.
[0318] Feedback integration: Explicit and implicit feedback from the user is integrated into threshold adjustments. The system identifies the user's satisfaction with responses and adjusts future judgment criteria accordingly.5. Temporal Analyzing ModuleThe Temporal Patterns of Emotional Changes are Analyzed:
[0319] Sequence processing: A long short-term memory network (LSTM) is used to process time-series features, to capture dynamic changes in emotional states. The memory mechanism of LSTM allows the system to “remember” long-term interaction history and identify emotional development trends.
[0320] Attention mechanism: Attention weights focus on key time points in a sequence to identify critical moments of emotional transition. For example, when a user suddenly changes from a calm state to excitement, the system may capture this transition point.
[0321] Trend analyzing: The emotional development trend (rising, stable, or declining) may be predicted to assist the device in making forward-looking adjustments. Trend information is used to generate more coherent and natural device response sequences.6. Processing on Missed Modal
[0322] When the system receives a plurality of modalities of data (such as voice, touch, and vision), a quality of each modal's data may be different from each other. For example, quality of voice signal may degrade in noisy environments, or visual signals may become unreliable due to insufficient lighting. This mechanism for processing missed modal may improve the robustness of system decisions by evaluating the reliability of each modal and dynamically adjusting its weight.Modal Reliability Evaluationr_m=f(SNR_m,variance_m,historical_performance_m)
[0323] In this step, the reliability score r_m for each modal m is determined, which is a comprehensive score configured to determine the reliability of the information provided by the current modal m. The higher the score, the more trustworthy the modal's information. The following factors may be fully evaluated by function f.
[0324] SNR_m: Signal-to-noise ratio, which is the ratio of valid information to noise in the signal. For example, the signal-to-noise ratio of the voice modal may reach over 30 dB in a quiet environment but may drop to as low as 5 dB in a noisy environment.
[0325] variance_m: The variance of internal features within the modal, and reflects the richness and consistency of the information. High variance may indicate rich information or severe noise.
[0326] history_performance_m: The performance record of the modal in historical interactions, recent accuracy rates may be determined in a sliding window.
[0327] The function f may be implemented as a weighted sum or a more complex nonlinear function as follows:r_m=α·normalized(SNR_m)+β·quality_score(variance_m)+γ·historical_performance_m
[0328] wherein SNR_m (Signal-to-Noise Ratio): Signal-to-noise ratio, which is a ratio of the strength of the useful signal to the strength of background noise. Generally, the higher the signal-to-noise ratio, the better the signal quality and the more reliable the data.
[0329] normalized( . . . ): Normalization processing. Since the signal-to-noise ratios of different modalities may have different dimensions or numerical ranges, normalization may unify them into a comparable range (e.g., between 0 and 1), so as to facilitate subsequent weighted summation.
[0330] variance_m: Variance of the mth modal data. Variance may show a volatility or uncertainty of the data. In some cases, smaller variance may indicate more stable and consistent data with higher quality; however, in other cases, an appropriate level of variance may be normal.
[0331] quality_score( . . . ): This is a function that provides a quality score based on the variance variance_m. The specific form of this function may depend on the application scenario. For example, smaller variance may correspond to a higher quality score, or there may be an optimal range of variance.
[0332] historical_performance_m: This refers to a performance of the modal in past tasks or data processing. For example, if a sensor has consistently been accurate in the past, its historical performance score will be relatively high.
[0333] α, β, and γ are weighting coefficients that are adjustable according to the application scenario.Dynamic Weight Allocationw_m=softmax(r_m)
[0334] The reliability scores of each modal are converted into a weight distribution by using the softmax function, so that the sum of all weights is 1:w_m=exp(r_m) / Σ_i exp(r_i)
[0335] Such conversion is conducted, so that:
[0336] modals with high reliability have greater weight.
[0337] All modal weights are within a range of 0 to 1, and sum thereof is 1.
[0338] When the reliability of a modal is significantly higher than others, its weight approaches 1.
[0339] For example, if the reliability scores for the tactile, voice, and visual modals are 0.8, 0.3, and 0.2 respectively, after the conversion of softmax, the tactile modal's weight is approximately 0.65, which is much higher than the others.Weighted Feature FusionX_fused=Σ(w_m·X_m)
[0340] At last, the system fuses features from each modal through weighted summation:
[0341] X_m is a feature vector of modal m.
[0342] w_m is a weight of the modal.
[0343] X_fused is a fused feature vector.
[0344] With the technical solution as stated above, when certain modal signals are weak or missing, the system may rely on information from other reliable modals more. For example, in a dark environment where the visual modal is almost ineffective (weight close to 0), the system will primarily rely on tactile and voice modals.
[0345] In embodiments, the system may determine the reliability of each modal in real-time and dynamically adjusts the fusion weights to adapt to changes in modal quality under different environmental conditions, so that the robustness and adaptability of the system in complex environments may be improved.System Training and Optimization1. Data Collection and Preparation
[0346] To train an effective multimodal emotion recognition model, it is necessary to collect multimodal data in specific scenarios.
[0347] Data sources: Real interaction data is collected in an ethically compliant manner, or simulated datasets are established based on professional knowledge. All data collection activities comply with strict ethical standards and privacy protection measures.
[0348] Data organization: A structured dataset containing multimodal information such as voice, touch, and vision is established, so that the data from different modalities may be synchronized in terms of time. Data annotations include emotion classes, emotion intensity, and temporal development trends.
[0349] Preprocessing on Data: Raw data is cleaned, denoised, and normalized to ensure training quality. It is necessary for voice data to reduce environmental noise, and for tactile data to filter out noise caused by sensor jitter.2. Strategy for Training Model
[0350] A two-stage training strategy is adopted to improve model performance.
[0351] Pre-training stage: Pre-training is conducted on large-scale general emotion datasets to learn capabilities for fundamental emotional representation. These general datasets include open multimodal emotion databases, such as the CASIA dataset (voice emotion) and the NLPCC dataset (text emotion).
[0352] Fine-tuning stage: Fine-tuning is conducted on datasets specific to the target scenario to adapt to the emotional expression patterns of interactions in specific adult product scenarios. During fine-tuning, a smaller learning rate is used to retain the general knowledge obtained during pre-training while adapting to the characteristics of the specific scenario.
[0353] Multi-task learning: A plurality of tasks related to each other for training on e.g., emotion classification, confidence estimation, and intensity prediction may be conducted simultaneously to improve the model's overall performance. Multi-task learning enables the model to better understand the multidimensional nature of emotions.3. Loss Function
[0354] A composite loss function is configured to optimize multiple objectives.
[0355] Emotion classification loss: Cross-entropy loss is used to optimize the accuracy in recognizing emotion class. For imbalanced distribution of emotion class, weighted cross-entropy is adopted to ensure good performance for minority classes.
[0356] Confidence estimation loss: Mean squared error loss is used to optimize the accuracy in predicting confidence, so that the model outputs low confidence in uncertain situations to avoid incorrect decisions.
[0357] Intensity estimation loss: Mean squared error loss is used to optimize the accuracy in predicting emotion intensity. Intensity prediction is crucial for controlling the amplitude of device responses.
[0358] Temporal trend loss: Cross-entropy loss is used to optimize the accuracy in classifying emotion change trends.
[0359] Weight balancing: Adjustable weight coefficients are used to balance the contributions of different loss components, to form a final multi-task loss function. The weight coefficients may be adjusted according to the requirements of specific application scenarios.4. Model Evaluation and Optimization
[0360] The model's performance is comprehensively evaluated by using multiple metrics.
[0361] Classification accuracy: Evaluation is conducted on overall accuracy in classifying emotions.
[0362] Class F1 score: Evaluation is conducted on the balance between precision and recall rate for each emotion class.
[0363] Confidence calibration: Evaluation is conducted on a consistency between confidence prediction and actual accuracy.
[0364] Temporal consistency: Evaluation is conducted on a smoothness and coherence of emotion predictions over time.Model Optimization is Conducted Based on Evaluation Results:
[0365] Hyperparameter tuning: Hyperparameters such as network structure, learning rate, and regularization strength may be optimized.
[0366] Model compression: Model size may be reduced by using techniques such as knowledge distillation and quantization to meet the resource constraints of embedded devices.
[0367] Feature engineering: Analyzing is conducted on importance of different modal features so as to determine the most effective feature combinations.Control Signal Generation Module1. Action Parameter Vector Space
[0368] The action parameter vector A defines various physical dimensions controllable by the device, to form a multidimensional control space. These parameters work together to create coordinated responses that match emotional states.A=[a1,a2, . . . ,an]Each Component Represents an Independent Physical Control Dimension:
[0369] Vibration frequency (15-200 Hz): the oscillation rate of the vibration unit. Low frequencies (15-40 Hz) mean a deep sensation, while high frequencies (100-200 Hz) means sharp stimulation.
[0370] Vibration intensity (0-100%): the vibration amplitude, which may be controlled to affect perceived intensity.
[0371] Contraction frequency (0-60 times / min): periodic motion frequency of the contraction unit.
[0372] Contraction intensity (0-100%): pressure level of the contraction unit.
[0373] Swing amplitude (0-100%): angular range of device swinging.
[0374] Swing speed (0-100 the angular velocity of device swinging.
[0375] Temperature variation (±5° C.): deviation of surface temperature relative to the baseline temperature.
[0376] Tactile feedback (0-100%): pressure intensity of the tactile feedback unit.
[0377] There is physical limit and valid range for each parameter for safety. The system automatically ensures that all parameters remain within these limits while optimizing the perceived experience.2. Emotion-Action Mapping Algorithm2.1 Mapping Function Based on Emotion Vectors
[0378] The emotion vector E contains intensity values for multiple emotional dimensions, such as:E=[e1,e2, . . . ,em]
[0379] For example: E=[e_arousal, e_pleasure, e_dominance, e_intimacy, e_excitement]
[0380] wherein, e_arousal, e_pleasure, e_dominance, e_intimacy, and e_excitement represent intensities of various emotional dimensions, typically within a range [0,1].
[0381] The emotion vector is converted into action parameters by using a mapping matrix M:A=f(M·E+b)
[0382] Where: M is an m×n mapping matrix that defines the relationship between emotions and actions.
[0383] b is the baseline action bias vector, defining default behavior when there is no significant emotion input.
[0384] f is a nonlinear activation function, configured to limit the output to be within a reasonable range by using functions of sigmoid or tanh
[0385] For example, if the system obtains an emotion vector E=[0.8, 0.6, 0.3, 0.7, 0.9] during detection, which indicates a state of high arousal, medium-high pleasure, low dominance, high intimacy, and high excitement, the action parameters A=[110 Hz, 75%, 35 times / min, 65%, 40%, 60%, +2.5° C., 50%] may be generated by calculation of mapping, to instruct the device to produce corresponding physical responses.2. Optimization of the Mapping Matrix
[0386] The mapping matrix M is not fixed but may be continuously optimized in many ways.
[0387] Initial setup: An initial emotion-physiological response relationships may be established based on psychological and physiological research. For example, excitement is associated with rapid heartbeat and increased body temperature, and thus may be related to higher vibration frequencies and temperatures.
[0388] User feedback optimization: User satisfaction feedback is collected to construct an optimization objective function:J(M)=−Σ(satisfaction_i)+λ∥M∥2
[0389] where satisfaction_i is a score of user satisfaction for the i-th interaction, and λ|M∥2 is a regularization term to prevent overfitting.
[0390] The mapping matrix is optimized by using gradient descent:M_new=M_old−η·∇J(M)
[0391] where η is a learning rate, typically set between 0.01 and 0.05.
[0392] ∇ is a vector differential operator representing the gradient.
[0393] Personalization adjustment: A user-specific vector P is introduced to adjust the mapping relationship:A=f(M·E·P+b)where P is a diagonal matrix representing the user's sensitivity to different emotional dimensions. For example, some users may be more sensitive to excitement, and thus P values may be set to be larger correspondingly.3. Temporal Dynamic Control Mechanism3.1 Smooth Transition Algorithm
[0394] To prevent discomfort caused by abrupt changes in action parameters, the system employs exponential moving average for smooth transitions:A_t=α·A_{t−1}+(1−α)·A_target
[0395] wherein A_t is an action parameter vector at time t.
[0396] A_{t−1} is an action parameter vector at the previous time.
[0397] A_target is an target action parameter determined based on the current emotion.
[0398] α is a smoothing coefficient that controls the transition speed.
[0399] The smoothing coefficient α is adjusted dynamically based on a type of transition.
[0400] For excited states requiring rapid responses, a is set to be smaller (0.3-0.5).
[0401] For relaxed states requiring gradual transitions, a is set to be larger (0.7-0.9).
[0402] For example, when user's emotion changes from a relaxed state (vibration frequency 30 Hz) to an excited state (target frequency 120 Hz), the system may be not changed sharply immediately but performs transition smoothly as follows (α=0.7):t=1: 30 Hz→57 Hz=0.7×30+0.3×120t=2: 57 Hz→75.9 Hz=0.7×57+0.3×120t=3:75.9 Hz→88.1 Hz
[0403] . . . gradually approaching the target value of 120 Hz.3.2 Rhythm Synchronization Mechanism
[0404] For periodic actions (e.g., user touches repeatedly or device vibrates), the system employs a phase-locked loop (PLL) algorithm to achieve synchronization:3.2.1 Frequency Determining:
[0405] The primary frequency f_user may be determined by performing analyzing on user's input signal by using fast Fourier transform (FFT) analysis.f_user=argmax(|FFT(input_signal)|)
[0406] This may refer to the raw data of user's actions received by the system. For instance, if the user repeatedly rubs the electronic skin with force, the input_signal may be a data sequence in which touch pressure or touch position changes over time. If the user shakes the device, the input_signal may be a time sequence of accelerometer readings.
[0407] argmax(IFFT(input_signal)|) may determine a frequency with the largest amplitude (intensity) among all frequency components.3.2.2 Phase Difference Calculation:Δφ=φ_user−φ_device
[0408] wherein φ user and φ_device are the phases of the user's actions and the device's actions, respectively.3.2.3 Frequency Adjustment:f_device′=f_device+K·Δφ
[0409] wherein K is a proportional coefficient, typically set to 0.1-0.3.
[0410] This mechanism ensures that the device's actions may synchronize with the user's actions in rhythm, to improve the coordination of interaction. For example, when the user is found to perform periodic touches at a frequency of 1.2 Hz, the device may automatically adjust its vibration frequency so that each vibration peak aligns with the user's touch actions, so that the user may have a more harmonious interactive experience.4. Examples of Action Parameters4.1 Controlling on VibrationAlgorithm for Controlling Frequencyf_vibration=f_base+β1·E_intensity+β2·E_arousal
[0411] f_vibration is a vibration frequency of the device. f_base is a preset default vibration frequency when there is no significant emotional input or the emotional state is relatively neutral. E_intensity represents the overall intensity of the user's detected emotion. This may be a composite value, e.g., in a range from 0 (no emotion) to 1 (very intense emotion). E_arousal is a dimension of emotion describing the physiological activation level of the emotion, with a range from calm and relaxed to excited and tense.
[0412] For example, anger and excitement are high-arousal emotions, while sadness and calmness are low-arousal emotions. This value is typically normalized within a range, such as [0,1] or [−1,1].
[0413] β1 and β2 are weight coefficients.
[0414] For example, when f_base=30 Hz, β1=50, and β2=70:
[0415] in a low-intensity, low-arousal state (E_intensity=0.2, E_arousal=0.1):
[0416] f_vibration=30+50×0.2+70×0.1=47 Hz (gentler vibration).
[0417] in a high-intensity, high-arousal state (E_intensity=0.8, E_arousal=0.9):
[0418] f_vibration=30+50×0.8+70×0.9=103 Hz (strongly excitant vibration).
[0419] The vibration frequency of the device is determined by a base frequency Lbase and is adjusted based on the current user's emotion intensity and arousal level. The stronger the emotion or the higher the arousal level (assuming β1 and β2 are positive), the vibration frequency may be increased accordingly, so that the vibration may be felt as being more “urgent” or “energetic.” This is a linear model that intuitively maps emotional features to vibration frequencies.Algorithm for Controlling Intensity:I_vibration=I_min+(I_max−I_min)·(E_intensity / E_max){circumflex over ( )}γI_vibration is a vibration intensity of the device. I_min is a minimum vibration intensity, and is the lowest vibration intensity that the device may generate or is configured to allow. Even emotion intensity is quite low, the vibration intensity will not fall below this value. I_max is the maximum vibration intensity the device may produce or is configured to allow. E_intensity represents the overall intensity of the user's emotion obtained by detection. E_max is an preset maximum value of emotional intensity configured to normalize the current emotional intensity E_intensity.
[0420] γ is a nonlinear exponent adjustment parameter configured to adjust the nonlinear relationship between emotional intensity and vibration intensity.
[0421] If γ=1, the mapping is linear: the vibration intensity increases in proportion to the increase in emotion intensity within the range of I_min and I_max.
[0422] If γ>1, the mapping curve bends upward (convex function): vibration intensity grows slowly at low emotion intensity and increases sharply at high emotion intensity.
[0423] This makes vibrations weaker for low-intensity emotions and amplifies vibrations for high-intensity emotions.
[0424] If κ<γ<1, the mapping curve bends downward (concave function): vibration intensity grows faster at low emotional intensity and slows down at high emotional intensity, so that noticeable vibrations may be generated for even low-intensity emotions, while the growth of vibration intensity for high-intensity emotions is less dramatic.
[0425] When I_min=10%, I_max=100%, E_max=1.0, and γ=2.0:
[0426] in moderate emotional intensity (E_intensity=0.5):
[0427] I_vibration=10+(100−10)×(0.5 / 1.0){circumflex over ( )}2.0=32.5%.
[0428] in high emotional intensity (E_intensity=0.9):
[0429] I_vibration=10+(100−10)×(0.9 / 1.0){circumflex over ( )}2.0=81.1%.4.2 Controlling on Humanoid Bionic MotionControlling on Facial Expression:
[0430] Facial motion parameter vectors are determined based on the emotion vector:F=[f_eye_blink,f_mouth_open,f face_flush, . . . ]ExamplesEye Blinking Frequency:f_eye_blink=max(0.5,min(5.0,2.0+3.0·E_arousal))
[0431] E_arousal is a dimension of emotion.
[0432] When E_arousal=0.2, f_eye_blink=2.6 times / minute (relaxed state).
[0433] When E_arousal=0.8, f_eye_blink=4.4 times / minute (excited state).Mouth Opening Degree:f_mouth_open=sigmoid(0.8·E_arousal+0.5·E_pleasure−0.3)
[0434] The sigmoid function (also known as the logistic function) is usually expressed as S(x)=1 / (1+e{circumflex over ( )}(−x)), where e is the base of the natural logarithm.
[0435] When E_arousal=0.7 and E_pleasure=0.6:
[0436] f_mouth_open=sigmoid(0.8×0.7+0.5×0.6−0.3)=sigmoid(0.56+0.3−0.3)=sigmoid(0.56)≈0.64, and in such case, the mouth is in a moderately open state.Coordination Control for Limb Motion:
[0437] A coupled oscillator network is used to realize coordinated motion of multiple joints:θ_i(t)=A_i·sin(ω_i·t+φ_i)+Σj K_ij sin(θ_j(t)−θ_i(t)−ψ_ij)
[0438] θ_i(t) is a phase (or angle) of the i-th oscillator at time t. θ_j(t) is a phase (or angle) of the j-th oscillator at time t. ψ_ij is a constant parameter that defines the phase relationship between oscillator i and oscillator j as desired.
[0439] A natural and coordinated limb motion patterns may be generated by using this algorithm.
[0440] For example, when a state “gentle stroking” is simulated, the arm and wrist joints move in coordination at a low frequency (ω≈0.5 Hz), with a moderate amplitude (A≈30°), and at a specific phase difference (φ≈90°).
[0441] When a state “enthusiastic hugging” is simulated, multiple joints coordinate at a medium frequency (ω≈1.2 Hz) and at a large amplitude (A≈60°).
[0442] The system may generate natural, smooth motions that conform to the biomechanical characteristics of the human body by adjusting the amplitude A_i, frequency ω_i, phase φ_i, and coupling coefficient K_ij.Optimization for Specific Scenarios1. Low Signal-to-Noise Ratio EnvironmentOptimization for Weak Sound Recognition:
[0443] Preprocessing for spectral enhancement: Convolutional filters are used to enhance the audio spectrum to strengthen signals in key frequency bands while suppressing noise interference. The system specifically focuses on frequency bands primarily associated with human voices (e.g., 300 Hz-3 kHz) to enhance signals in these regions.
[0444] Extraction of self-supervised feature: Self-supervised learning based on large-scale audio data is employed to improve the model's sensitivity to weak signals. Self-supervised learning does not rely on manual annotations and may leverage large amounts of unlabeled audio data to learn effective sound representations.
[0445] Attention mechanism enhancement: particular attention mechanisms are adopted to automatically focus on key features of weak audio signals. The system may “listen” to subtle sound variations below 20 dB, even in a noisy environment.
[0446] Multi-scale feature fusion: Acoustic features at different time scales are combined to improve the recognition of weak non-verbal sounds (e.g., moans, sighs). The system conducts analyzing on both short-term features (e.g., instantaneous pitch) and long-term patterns (e.g., rhythm and emotional development).2. High-Precision Action RecognitionOptimization for Fine Motion Detection:
[0447] Processing based on sliding window: Sliding window techniques are applied to analyze continuous motion data to accumulat micro-motion changes within 30 seconds. Analyzing based on window allows the system to capture gradual motion changes rather than just instantaneous states.
[0448] Bidirectional GRU temporal modeling: Bidirectional gated recurrent unit (GRU) networks are configured to process sequential data, to capture motion rhythms and trends. The gating mechanism of GRU learns both short-term and long-term motion patterns effectively.
[0449] Trend detector: A particular neural network branch may be used to conduct analyzing on changes in motion speed and amplitude. The system may identify motion trends such as “acceleration,”“constant speed,” or “deceleration,” to provide predictive guidance for device responses.
[0450] High-precision sensor algorithms: A particular filtering and enhancement algorithm is used for millimeter-wave sensors and pressure sensors to detect acceleration changes as low as 0.2 mm / s2 and pressure variations within the range of 1-5 g. Sensor data is processed by multi-stage low-pass filtering to remove noise while retaining true motion signals.3. Confidence-Weighted Fusion MechanismOptimization for Modal Imbalance:
[0451] Modal-specific confidence: Reliability of data of each modal is independently evaluated to generate modal confidence scores. For example, in a noisy environment, the confidence of voice signals may be lower; when touch sensors are affected by external vibrations, the confidence of tactile signals may decrease.
[0452] Dynamic weight allocation: The weight of each modal in the fusion process is adjusted based on their confidence levels dynamically. When a quality of signal of a specific modal decreases, the system reduces its influence weight and prioritizes modal signals with higher-quality automatically.
[0453] Adaptive fusion strategy: The system selects the most suitable modal combinations and fusion ways automatically based on different scenarios and user habits. The system learns each user's interaction pattern preferences and adjusts the modal weight allocation strategy accordingly.EXAMPLES OF PRACTICAL APPLICATIONS1. Example 1: Interaction Between a User and an Intelligent Silicone Doll
[0454] Processing may be as follows:
[0455] 1.1 Collecting multimodal data: A voice recognition system captures and transcribes the speech of “You look beautiful today.”
[0456] a. Simultaneously extracts acoustic features of the voice (gentle tone, moderate speed).
[0457] b. The pressure sensor array detects gentle touch pressure (approximately 5-10 g).
[0458] c. The capacitive sensor tracks touch trajectories (slow and gentle facial touch).
[0459] 1.2 Feature extraction: The language model extracts semantic features from the text and identifies the semantics of praise and intimacy.
[0460] a. The acoustic analyzer determines the tone is gentle and the emotion is positive.
[0461] b. The tactile analyzer identifies the touch pattern as “gentle stroking,” expressing an intimate emotion.
[0462] 1.3 Multimodal fusion: The Transformer fusion layer performs cross-attention processing on linguistic, acoustic, and tactile features.
[0463] a. Captures the consistency between the speech content and the touch pattern (both expressing gentle emotions).
[0464] b. Generates a unified multimodal representation, reinforcing the emotional information conveyed by all modalities.
[0465] 1.4 Emotion classification: The system identifies the emotion class as “gentle / intimate” with 85% confidence.
[0466] a. The emotional intensity is evaluated as moderately high (0.75 / 1.0).
[0467] b. Trend analysis shows the emotion is stable with a slight upward trend.
[0468] 1.5 Generating response by devices: Based on the emotion class “gentle / intimate”, the corresponding device response mode is selected.
[0469] a. The facial expression controller is used to generate a smiling expression and slow blinking action.
[0470] b. Facial heating elements are activated to simulate a shy blushing reaction (temperature rises by 2° C.).
[0471] c. A gentle voice response is generated: “Thank you, I think you look particularly charming today too”.
[0472] d. The limb controllers are coordinated to perform a gentle cuddling motion.
[0473] 1.6 Feedback and adaptation: Interaction data is recorded to update the user preference model.
[0474] a. the baseline thresholds are adjusted for future interactions, so that the system is more sensitive to gentle interactions with this user.2. Example 2: Interaction Between a User and an Intelligent Massager
[0475] Processing may be as follows.
[0476] 2.1 Collecting Data in a low signal-to-noise ratio environment: An enhanced acoustic processor captures weak speech or non-verbal sounds, such as “Can you go a bit faster?”
[0477] a. Spectral enhancement technology may be used to strengthen key frequency bands, to improve signal quality.
[0478] b. High-precision accelerometers may be used to capture subtle wrist acceleration movements.
[0479] c. The temporal processor accumulates motion variation data over 30 seconds.
[0480] 2.2 Enhanced feature extraction: An acoustic analyzer extracts features of speech / non-verbal sounds for weak signals.
[0481] a. A motion analysis system identifies an increasing trend in wrist motion frequency.
[0482] b. The pressure analyzer determines pressure variation patterns, to determine changes in intensity during using.
[0483] 2.3 Confidence-weighted fusion: Confidence scores for voice and motion data are determined.
[0484] a. In a quiet environment, the confidence of voice data may be low (approximately 40%).
[0485] b. Motion data confidence is relatively high (approximately 90%).
[0486] c. The system increases the weight of motion data and reduces the weight of voice data automatically.
[0487] d. A fused representation is generated, primarily based on motion features supplemented by voice features.
[0488] 2.4 State recognition: The system identifies the state as “excited / anticipating reinforcement” with 93% confidence.
[0489] a. Intensity evaluation shows a gradual increase (from 0.60 to 0.85).
[0490] b. Trend analysis indicates an accelerating trend, expected to continue strengthening.
[0491] 2.5 Adjusting device parameter: Based on the identified state and trend, a plan for adjusting device parameter is generated.
[0492] a. The vibration frequency gradually increases from 65 Hz to 90 Hz (initial adjustment completed within 3 seconds).
[0493] b. The amplitude intensity is gradually increased from 60% to 75%.
[0494] c. The vibration mode transitions to a wave-like fluctuation pattern, to increase variation.
[0495] d. The device temperature slightly rises (approximately 2° C.), to increase tactile stimulation.
[0496] Continuous monitoring and fine-tuning: The system continuously monitors user responses and optimizes parameters every 20 seconds.
[0497] a. When changes in user voice features are detected, further adjustments are made to the vibration mode.
[0498] b. The system records the user's reactions to specific parameter combinations to optimize personalized configurations.Privacy Protection and Security ProcessingData Security
[0499] Local processing: Analyzing on core emotion is conducted locally on the device, and sensitive data is not uploaded to the cloud.
[0500] Data anonymization: Necessary data collection undergoes strict anonymization processes to remove personally identifiable information.
[0501] Differential privacy: Differential privacy technology is applied to protect user preference data by adding noise, which is well-controlled.
[0502] Data encryption: All stored and transmitted data is protected with highly complicated encryption.System Security
[0503] Secure boot: Devices use secure boot mechanisms to ensure only authorized software is executed.
[0504] Access control: Strict permission management is implemented to restrict access to sensitive functions.
[0505] Regular updates: Security updates are provided regularly to patch potential vulnerabilities.
[0506] User authorization: All advanced functions require explicit user authorization to be enabled.Deployment and OptimizationModel Compression and Optimization
[0507] Model quantization: Floating-point operations are converted to fixed-point operations to reduce model size and accelerate inference.
[0508] Knowledge distillation: Core knowledge is extracted from complex models to be used in training for a lighter model.
[0509] Architecture optimization: Network architecture is optimized for embedded devices to reduce complexity in computing.
[0510] Selective activation: some of components in a model are selectively activated based on scenario requirements to save energy.Embedded Deployment
[0511] Resource allocation: Computation resources are allocated as needed based on the device's hardware capabilities.
[0512] Power management: Intelligent power management strategies are implemented to extend battery's lifetime.
[0513] Thermal management: Thermal management is optimized to ensure stability during operation in long time.
[0514] Asynchronous processing: Asynchronous processing mechanisms are adopted to ensure smooth user interface responsiveness.Summary
[0515] The system is configured as a complete multimodal emotion recognition and interaction solution, which may be embedded into adult products. The system may provide high-precision emotion recognition by using the following key features:
[0516] 1. Multimodal feature extraction and fusion to process linguistic, acoustic, tactile, and visual data.
[0517] 2. Cross-modal attention mechanism based on Transformer configured to capture correlations among various modalities.
[0518] 3. Temporal analysis module configured to perceive emotion changing trends.
[0519] 4. Adaptive threshold mechanism configured to dynamically adjust confidence scores based on historical interactions.
[0520] 5. Two-stage training strategy for pretraining based on general data, which is followed by fine-tuning for specific scenarios.
[0521] The system generates corresponding device responses based on recognized emotion states, to provide a more natural and immersive interactive experience while ensuring user data privacy and system security.Step 1: Feature Extraction
[0522] When multimodal information is input from user, such as voice, image, text, body motion, or external environmental information involving two or more modalities, features of each single modal are extracted separately to obtain corresponding feature vectors, as follows:(1) Voice Modal Features:
[0523] The voice emotion analysis module is used to extract emotional features from the voice, such as fundamental frequency, pitch, and speech rate.
[0524] Example of feature vector: F_voice=[0.5, 0.8, 0.3].(2) Image Modal Features:
[0525] The image processing module is used to extract visual features from user expressions or movements, such as facial key points and motion postures.
[0526] Example of feature vector: F_image=[0.2, 0.6, 0.4].(3) Text Modal Features:
[0527] The natural language processing module (e.g., pretrained language model BERT) is used to extract semantic features.
[0528] Example of feature vector: F_text=[0.9, 0.7, 0.1].(4) Body Motion Modal Features:
[0529] The motion capture module (e.g., inertial measurement units or cameras) is used to extract the user's motion trajectories and frequency features.
[0530] Example of feature vector: F_motion=[0.4, 0.5, 0.7].(5) External Environmental Modal Features:
[0531] Sensors are used to collect environmental data (e.g., brightness, temperature, sound) and extract features.
[0532] Example of feature vector: F_env=[0.3, 0.2, 0.6].
[0533] It should be understood that in other embodiments, external environmental modal features may be unused.Step 2: Feature Normalization
[0534] Each modal's feature vector is normalized to bring all feature values into the same range (e.g., [0, 1]) to balance influences of various modalities on subsequent fusion.
[0535] For example, after normalization, F_voice=[0.5, 0.8, 0.3] may become F_voice_normalized=[0.625, 1.0, 0.375].Step 3: Concatenation Fusion
[0536] The normalized features of all modalities are directly concatenated into a high-dimensional feature vector, to form a preliminary fused feature representation. The concatenated feature vector contains complete information from all modalities, as shown below:
[0537] For example:
[0538] Voice modal features: F_voice_normalized=[0.625, 1.0, 0.375].
[0539] Image modal features: F_image_normalized=[0.2, 0.6, 0.4].
[0540] Text modal features: F_text_normalized=[0.9, 0.7, 0.1].
[0541] Body motion modal features: F_motion_normalized=[0.4, 0.5, 0.7].
[0542] After concatenation,
[0543] F_concat=[0.625, 1.0, 0.375, 0.2, 0.6, 0.4, 0.9, 0.7, 0.1, 0.4, 0.5, 0.7].Step 4: Modeling Inter-Modal Relationships (Based on Attention Mechanism)
[0544] An attention mechanism is applied on the concatenated high-dimensional feature vector to determine a weight of each modal to highlight the most important modal features for the current scenario. Specifically, this step 4 may be conducted as follows:(1) Determining Attention Weights:
[0545] A self-attention mechanism or multi-head attention mechanism is used to adjust modal weights dynamically based on contextual information (e.g., the user's current emotional state).
[0546] Example of weights:
[0547] Voice modal weight: α_voice=0.4.
[0548] Image modal weight: α_image=0.3.
[0549] Text modal weight: α_text=0.2.
[0550] Body motion modal weight: α_motion=0.1.(2) Weighted Feature Fusion:
[0551] The concatenated features are weighted by modal weights to obtain the weighted fused feature representation:F_weighted=α_voice*F_voice_normalized+α_image*F_image_normalized+α_text*F_text_normalized+α_motion*F_motion_normalized.
[0552] F_weighted is a final fused feature vector. It integrates information from voice, image, text, and motion modalities. This F_weighted may be used as input for subsequent tasks such as emotion recognition, behavior prediction, or decision-making.
[0553] α_-voice is a weight of the voice modal, and F_voice_normalized is the normalized voice feature vector.
[0554] α_image*F_image_normalized is a weighted normalized image feature.
[0555] α_text*F_text_normalized is a weighted normalized text feature.
[0556] α_motion*F_motion_normalized is a weighted normalized motion feature.Step 5: Deep Feature Extraction
[0557] The weighted fused features are input into a deep learning network (e.g., a multilayer perceptron MLP or Transformer) to further extract nonlinear relationships between modalities and generate the final unified feature representation.(1) Deep Learning Network:
[0558] Input layer configured to receive fused feature vector F_weighted=[0.53, 0.77, 0.36].
[0559] Hidden layers configured to extract complex inter-modal interactions through a multilayer perceptron (MLP) or Transformer.
[0560] Activation function configured to conduct activation by using ReLU or GELU to enhance the nonlinear representation capability of features.
[0561] Output layer configured to generate the final unified feature representation.(2) Final Feature Representation:
[0562] Example of output vector:
[0563] F_final=[0.6, 0.8, 0.5, 0.4, 0.7].
[0564] This vector contains the integrated semantic features of multimodal information and can be directly input into a multimodal model (e.g., GPT) to be used in recognizing user's intent.
[0565] It should be understood that the above descriptions of features and their vectors are merely examples provided to facilitate understanding by skilled in the art. In other embodiments, other vector representations may be used, and the present disclosure is not limited by the above examples.
[0566] Further, as mentioned in the above embodiment, the emotion state quantization feature is an indicator for expressing the user's emotions quantitatively so as to reflect the changes in the user's emotions accurately during the use of adult products. With the emotion state quantization feature, precise control or dynamic adjustment may be applied on the operating parameters of adult products (such as vibration intensity, frequency, motion mode, movement amplitude, etc.). Therefore, to facilitate the understanding of the specific implementation of the emotion state quantization feature in the control process, an exemplary description is provided below. Reference could be made to the above introduction for any operation, which is not described in detail herein. Specifically, this example includes the following.1. Range and Definition of Emotion State Quantization Features
[0567] In this example, the value range of the emotion state quantization feature is set to 0-10, with specific classifications and meanings as follows:
[0568] 0: No emotion detected (noise or invalid data).
[0569] 1-2: Mild state (e.g., slight liking or low excitement).
[0570] 3-5: Normal state (e.g., moderate liking or medium excitement).
[0571] 6-7: Strong state (e.g., strong liking or high excitement).
[0572] 8-10: Extreme state (e.g., extreme liking or very high excitement).2. Examples of Emotion State Quantization Features Set in Applications
[0573] During control of adult products, the emotion state quantization feature may be directly used to adjust the operating parameters of the product. For example, when the emotional feature is “liking,” vibration intensity and frequency may be set based on the quantization feature value correspondingly.i1. Mild State (1-2):
[0574] Vibration intensity is set to the lowest range (e.g., 20%-40% of power).
[0575] Vibration frequency is maintained at a low range (e.g., 1-3 Hz).i2. Normal State (3-5):
[0576] Vibration intensity is set to a medium range (e.g., 50%-70% of power).
[0577] Vibration frequency is maintained at a medium range (e.g., 4-6 Hz).i3. Strong State (6-7):
[0578] Vibration intensity is set to a higher range (e.g., 80%-90% of power).
[0579] Vibration frequency is increased to a higher range (e.g., 7-9 Hz).i4. Extreme State (8-10):
[0580] Vibration intensity is set to the maximum safe range (e.g., 95%-100% of power).
[0581] Vibration frequency reaches the highest allowable frequency of the device (e.g., 10-12 Hz).
[0582] It should be noted that the above parameter ranges are within the safe design range of adult products and may be preset or adjusted based on user preferences. In other embodiments, there may be more or fewer states set for emotion features. Generally, one setting for emotion states may includes at least two or more quantized classifications. The more states are set, the higher the precision requirement for user intent recognition, and correspondingly, the finer the control of the adult product. By using the emotion state quantization feature for multi-level classification of the same emotional state, the requirements on intelligent interaction between the user and the adult product during use may be satisfied more, and the emotion interaction with the user becomes more precise.3. Generation of Emotional State Quantization Features
[0583] In this example, the emotional state quantization feature is generated by using multimodal emotion recognition. The operational steps may be as follows:(1) Data Collection
[0584] Emotion-related data is collected by using multimodal information, including but not limited to voice data, image data, and external environment data, wherein:
[0585] Voice data is obtained by collecting user's voice through a microphone and extracting voice features such as pitch, volume, and speech rate.
[0586] Image data is obtained by collecting user's facial expression or body movements through a camera and extracting visual features such as facial key points or body movements.
[0587] External environment data is obtained by collecting environmental information through sensors to facilitate emotion recognition and the environmental information may be e.g., background music and light intensity.
[0588] In other examples, one or more of the above data may be collected, e.g., external environment data may be omitted.(2) Feature Extraction
[0589] Features are extracted from the collected data to generate multimodal emotional features:i1. Voice Emotion Features, Including:Extracting features such as fundamental frequency, volume, speech rate, and pitch from audio signals.
[0591] Generating emotion vectors.i2. Facial Expression Features:
[0592] Determining facial key points (e.g., positions of eyebrows, mouth corners, and eyes) and extracting expression features.
[0593] Using a facial expression recognition model to predict emotion classes.i3. Body Motion Features:
[0594] Analyzing the user's motion trajectory, frequency, and intensity, such as through an inertial measurement unit (IMU).
[0595] Using a motion pattern classifier to identify motion features related to emotions.(3) Modal Fusion
[0596] Multimodal features are fused into a unified emotional feature vector, specifically including the following operations:
[0597] Concatenation fusion: Directly concatenating features of each modal into a high-dimensional feature vector.
[0598] Weighted summation fusion: Adjusting modal weights dynamically by using an attention mechanism to highlight key modalities.
[0599] Deep fusion: Using a fusion network (such as Transformer or a multilayer perceptron MLP) to extract nonlinear relationships between modalities and generate a unified emotional vector.4. Emotional State Quantization
[0600] An emotion classifier is added to the neural network model to quantify the emotional state. The specific operations are as follows:(1) Classifier:
[0601] Using a fully connected layer with a Softmax activation function to classify emotion classes.
[0602] Example of output class: liking, disliking, excitement, calmness, etc.(2) Quantization Feature Generation:
[0603] Quantized feature values (in a range of 0-10) are generated based on the predicted scores of emotion class in connection with contextual information.Example
[0604] If the emotion is classified into “liking” with a predicted score of 0.85, a quantized value of 7.5 is generated according to the mapping rule.
[0605] Of course, the mapping rule may be adjusted based on experimental data. The above is only an example.Real-Time Updating and Feedback
[0606] During the use of adult products, real-time emotion recognition and quantization adjust the operating state of the product dynamically.
[0607] For example, the system updates the emotional state quantization feature in a preset time interval (e.g., 5 seconds) and then adjusts parameters (such as motion speed, vibration intensity, and frequency) based on the latest quantized feature value.
[0608] With the above example process, the implementation of emotion state quantization in the embodiments of the present disclosure may be understood clearly. Of course, in other embodiments, other methods for setting emotional state quantization may be adopted.
[0609] In the above embodiments of the present disclosure, the system for interactive controlling on adult product based on a multimodal model may be implemented by referring to any of the technical steps in the above-mentioned methods.
[0610] Based on same technical concept of the method for interactive controlling on adult product based on a multimodal model, another embodiment of the present disclosure further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The processor executes the program to implement the steps of the method for interactive controlling on adult product based on a multimodal model.
[0611] Refer to FIG. 4, which shows a block diagram of an electronic device according to an exemplary embodiment. As shown in FIG. 4, the electronic device may include: processor 401, memory 402. The electronic device may further include one or more of multimedia component 403, input / output interface 404, communication component 405, and power components 406. Processor 401 obtains user-side information and extracts feature representations, identifies user intent based on the multimodal model, generates control commands based on user intent, and converts them into control signals, which are sent to the control system of the adult product through communication component 405 to execute corresponding actions.
[0612] Processor 401 is configured to control overall operation of the electronic device to complete all or part of the steps of the method or interactive controlling on adult product based on a multimodal model in the first aspect. Memory 402 is configured to store various types of data to support the operation of the electronic device. Such data may include instructions for any application or method configured to operate on the electronic device, as well as data related to applications, such as user's information, received and sent messages, pictures, audio, video, etc.
[0613] Multimedia component 403 may include a camera and audio component 410, etc., and may further include a screen as needed. The screen may be a touch screen, and the camera may be configured to capture images. Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 may include a microphone configured to receive external audio signals. The received audio signals may be further stored in memory 402 or sent through communication component 405. Audio component 410 includes at least one speaker configured to output audio signals.
[0614] Input / output interface 404 provides an interface between processor 401 and other interface modules, which may include a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 405 is configured to enable wired or wireless communication between the electronic device and other devices. Wireless communication may include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G technologies, or a combination of one or more of them, which is not limited herein. Therefore, the corresponding communication component 405 may include: Wi-Fi module, Bluetooth module, NFC module, etc.
[0615] The electronic device in this embodiment can be provided on adult products or set independently. For example, it can be made into a very small control box provided in the head position of an adult product doll, similar to the existence of a human brain. It forms control commands based on the acquired user-side information and user intent to control the use of the adult product doll. It can also be made into a small component set within the adult product.
[0616] Alternatively, the memory is configured to store programs. The memory may include volatile memory, such as random-access memory (RAM), static random-access memory (SRAM), or Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc. The memory may further include non-volatile memory, such as flash memory. The memory is configured to store computer programs (such as applications or functional modules that implement the above methods), computer instructions, etc. The computer programs, computer instructions, etc., may be partitioned and stored in one or more memories. Furthermore, the computer programs, computer instructions, and data may be called by the processor.
[0617] The processor is configured to execute the computer program stored in the memory to implement the steps of the methods described in the above embodiments. Specific implementation may refer to the relevant descriptions in the previous method embodiments.
[0618] The processor and memory may have an independent structure or be integrated into a single structure. When the processor and memory are independent structures, the memory and processor may be coupled and connected via a bus.
[0619] In addition, the embodiments of the present disclosure may further provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer program is executed by the processor, it implements the steps of the method for interactive controlling on adult product based on a multimodal model
[0620] The computer-readable medium includes computer storage media and communication media, where communication media include any medium that facilitates the transfer of computer programs from one place to another. The storage media may be any available medium accessible by a general-purpose or special-purpose computer. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium may also be part of the processor. The processor and storage medium may be located in an ASIC. Additionally, the ASIC may be located in user equipment. Of course, the processor and storage medium may also exist as separate components in communication devices.
[0621] Those skilled in the art can understand that all or part of the steps of the above-described method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments. The aforementioned storage medium includes ROM, RAM, magnetic disks, optical disks, or other media capable of storing program codes.
[0622] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, without limiting them. Although the present disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments or replace some or all of the technical features with equivalents. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present disclosure. The above preferred features can be used in any combination as long as they do not conflict.
Examples
embodiment 1
The user is male, and the adult product is a smart silicone doll
[0245]User behavior: The user says to the smart silicone doll, “You look beautiful today”, while touching the doll's cheek gently and slowly.
[0246]Emotion recognition: four input signals may be processed by the multimodal Transformer model: voice content (converted into text via speech recognition), voice acoustic features (such as fundamental frequency, energy, rhythm, etc.), touch pressure data (collected through a pressure sensor array), and touch speed / trajectory data (obtained in real-time by capacitive sensors). Each modal data is first processed by an independent feature extractor and then fused across modalities by using a self-attention mechanism to capture correlations between different signals. The model adopts a pretraining-finetuning paradigm to be pretrained on large-scale emotional interaction data and then finetuned for specific scenarios. Confidence scores may be dynamically adjusted based on historical...
embodiment 2
The User is Female, and the Adult Product is a Smart Massager
[0255]User behavior: During use, the user softly says, “Can you go faster,” or makes weak vocalizations such as moaning, while slightly accelerating the movement of the massager by her wrist.
[0256]Emotion recognition: The multimodal Transformer architecture employs a specially enhanced processing flow designed for low signal-to-noise ratio environments. The voice processing branch uses an acoustic model specifically trained for weak sounds and non-verbal vocalizations (such as moaning and panting) by self-supervised learning means to extract feature patterns from large-scale contextual audio. The model employs a specialized attention mechanism capable of capturing subtle sound variations below 20 dB in noisy environments, with spectral enhancement techniques to amplify critical frequency bands.
[0257]High-precision millimeter-wave sensor array capable of recognizing subtle acceleration changes as low as 0.2 mm / s2 may be use...
embodiment 3
The User is Male, and the Adult Product is a Smart Silicone Doll or a Device with a Similar Function
[0279]Pressure sensors are provided on chest and other locations of the smart silicone doll. Temperature sensors are provided on the genital area and other locations of the smart silicone doll. Other types of sensors, such as cameras, light sensors, or sound sensors, are provided in positions as needed on the smart silicone doll. A smart terminal may be used as a camera, light sensor, or sound sensor.
[0280]Rotary actuators and linear actuators are provided at each joint and movable part of the smart silicone doll. Each actuator is connected to a motor or driving component, and a clamping mechanism is provided in the genital area.
[0281]The main control component conducts analyzing on one or more types of collected perception data and determines control instructions according to the above method, so as to control one or more of eyes, mouth, head, shoulders, elbows, wrists, hips, knees, ...
Claims
1. A method for dynamically controlling a physical actuator of an adult product, comprising:acquiring perception data of a user interacting with the adult product by at least one microphone and at least one sensor, wherein the perception data of the user comprises one or more modal data of voice data, tactile data, and visual data; the at least one sensor comprises at least one pressure sensor;processing the perception data of the user by using a processor to run a multimodal neural network model, which comprises: conducting feature-extracting on the one or more modal data in the perception data to obtain modal feature vectors corresponding to the one or more modal data respectively; applying an attention mechanism to each modal feature vector to obtain modal weights by calculation and generate a weighted cross-modal feature representation that quantifies an emotional state of the user; mapping the weighted cross-modal feature representation to a specific emotion class of the user selected from a predefined set of emotion classes;conducting a search in a data structure of action parameter vectors to generate a multi-component action parameter vector, A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set of emotion classes as a response to the specific emotion class of the user; each component of the action parameter vector represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions comprising at least vibration frequency and contraction intensity; andconverting the numerical control values of the action parameter vector into drive signals and transmitting the drive signals to the physical actuator, so that the adult product performs physical actions defined by the numerical control values,wherein the conducting feature-extracting on the one or more modal data in the perception data to obtain modal feature vectors corresponding to the one or more modal data respectively comprises:when the perception data comprises voice data, converting voice data using a pre-trained language model into voice text and extracting semantic representation vectors, extracting fundamental audio features from the voice data using Mel-frequency cepstral coefficients, and extracting acoustic feature vectors from the fundamental audio features using spectrum analysis, wherein the acoustic feature vectors comprise fundamental frequency, energy, and harmonic-to-noise ratio;when the perception data comprises tactile data, processing tactile data by using a multilayer perceptron network to obtain tactile features vectors corresponding to tactile data, wherein the tactile feature vectors comprise one or more of touch pattern, touch force, touch rhythm, touch trajectory, touch area, and action frequency; andwhen the perception data comprises visual data, extracting facial expression features based on a facial recognition neural network, determining gazing duration and gazing frequency using a pupil tracking algorithm to obtain eye contact feature vectors, and extracting body posture feature vectors using a human pose estimation model.
2. The method for dynamically controlling a physical actuator of an adult product according to claim 1, further comprising:acquiring real-time feedback data from the user, wherein the feedback data includes behavior response; anddynamically adjusting the predefined action parameter vectors stored in the data structure based on the real-time feedback data.
3. The method for dynamically controlling a physical actuator of an adult product according to claim 1, wherein after determining the specific emotion class of the user from the predefined set of emotion classes, the method further comprises:evaluating confidence of the emotion class by using a neural network.
4. The method for dynamically controlling a physical actuator of an adult product according to claim 1, further comprising:forming temporal sequence features from the modal feature vectors corresponding to each type of modal data within a preset period and determining dynamic changes in emotion recognitions by processing the temporal sequence features using a long short-term memory network, to determine an emotional trend;determining a moment of change in emotion, by focusing on transition points in the temporal sequence features with attention weights.
5. The method for dynamically controlling a physical actuator of an adult product according to claim 1, further comprising:collecting user's interaction data and feedback information;updating parameters of the multimodal model dynamically using deep learning algorithms and reinforcement learning algorithms;adjusting control logic to meet user's personalized needs adaptively and optimizing determining on user intent and control strategies.
6. The method for dynamically controlling a physical actuator of an adult product according to claim 1, further comprising:downloading and installing a latest multimodal model parameters or control logic through a network connection or from a cloud platform to conduct an update process, wherein the update process uses digital signatures and verification mechanisms to ensure authenticity and integrity of content being updated.
7. A method for interactive controlling on an adult product based on a multimodal model, comprising:acquiring perception data of a user interacting with the adult product, wherein the perception data of the user comprises one or more modal data of voice data, tactile data, and visual data;conducting feature-extracting on the one or more modal data in the perception data to obtain modal feature vectors corresponding to the one or more modal data respectively;conducting feature-fusing on the extracted modal feature vectors to form cross-modal feature representation so as to determine an emotion class of the user from a predefined set of emotion classes;conducting a search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product; andconverting the control instruction into drive signals for the adult product, so that the adult product performs actions corresponding to the drive signals,wherein the conducting feature-extracting on the one or more modal data in the perception data to obtain modal feature vectors corresponding to the one or more modal data respectively comprises:when the perception data comprises voice data, converting voice data using a pre-trained language model into voice text and extracting semantic representation vectors, extracting fundamental audio features from the voice data using Mel-frequency cepstral coefficients, and extracting acoustic feature vectors from the fundamental audio features using spectrum analysis, wherein the acoustic feature vectors comprise fundamental frequency, energy, and harmonic-to-noise ratio;when the perception data comprises tactile data, processing tactile data by using a multilayer perceptron network to obtain tactile features vectors corresponding to tactile data, wherein the tactile feature vectors comprise one or more of touch pattern, touch force, touch rhythm, touch trajectory, touch area, and action frequency; andwhen the perception data comprises visual data, extracting facial expression features based on a facial recognition neural network, determining gazing duration and gazing frequency using a pupil tracking algorithm to obtain eye contact feature vectors, and extracting body posture feature vectors using a human pose estimation model.
8. The method according to claim 7, wherein the conducting a search in a data structure of action parameter vectors according to the emotion class of the user, to determine a control instruction for the adult product comprises:conducting the search in a data structure of action parameter vectors to generate a multi-component action parameter vector A=[a1, a2, . . . , an], wherein the data structure of action parameter vectors stores predefined action parameter vectors for each emotion class in the predefined set as a response to the emotion class of the user; the control instructions are configured to adjust one or more of a motion mode, intensity, or frequency of the adult product,wherein, each component of A, a1, a2, . . . , an represents a specific numerical control value for an independent physical control dimension of the adult product, the physical control dimensions comprise one or more of vibration frequency, vibration intensity, contraction frequency, contraction intensity, swing amplitude, swing speed, temperature change amplitude, and pressure intensity of tactile feedback.
9. The method according to claim 7, wherein the conducting feature-fusing on the extracted modal feature vectors to form cross-modal feature representation so as to determine the emotion class of the user from the predefined set of emotion classes comprises:projecting each modal feature vector into a common dimensional feature space by linear mapping, to generate aligned modal features;determining attention scores between the aligned modal features by calculation, to determine modal weights for generating a weighted cross-modal feature representation, and generate the weighted cross-modal feature representation that quantifies an emotional state of the user;mapping the weighted cross-modal feature representation to the emotion class of the user selected from a predefined set of emotion classes.
10. The method according to claim 7, wherein the control instruction further comprises one or more of:voice control instruction, configured to make a selection among preset voice feedback templates and adjust voice parameters according to the emotion class of the user;atmosphere control instruction, configured to change one or more of brightness of a light source, color of the light source, flicker frequency of the light source, and background music, based on the emotion class of the user.
11. The method according to claim 7, further comprising:performing multi-factor authentication and authorization management on user's identity to ensure that only authorized users control current adult product;encrypting user data in a way of end-to-end to prevent unauthorized access and data leaks;monitoring an execution process of the control instructions to block abnormal instructions when being detected, so as to ensure safe use.
12. The method according to claim 7, further comprising:downloading and installing a latest multimodal model parameters or control logic through a network connection or from a cloud platform to conduct an update process, wherein the update process uses digital signatures and verification mechanisms to ensure authenticity and integrity of content being updated.
Citation Information
Patent Citations
OTA updating method and related device
CN117579606A
Video detection method and adult product
CN119131663A
Systems and methods for providing adaptive biofeedback measurement and stimulation
US20150305971A1
Light-Emitting Diode and Massage Device for Delivering Focused Light for Vaginal Rejuvenation
US20170095398A1
Male and female sexual aid with wireless capabilities
US20180140502A1