Device control method and apparatus, electronic device, and storage medium
By using multimodal input information and hierarchical verification methods, combined with a security encryption module, the problem of insufficient data transmission security in vehicle-home interconnection is solved, achieving higher security and privacy protection, and improving data processing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING CHANGAN AUTOMOBILE CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-17
AI Technical Summary
The lack of effective security encryption mechanisms in existing vehicle-to-home connectivity technologies leads to a high risk of user privacy leaks during data transmission.
Multimodal input information is used for device control. User control permissions are verified through hierarchical verification. The appropriate verification method is selected according to the privacy category of the target device to ensure that the user's control permissions match the device's privacy category. A security encryption module is introduced to ensure data transmission security.
It improves the security and privacy protection of vehicle-home connectivity, reduces the risk of user privacy leaks, and optimizes data processing efficiency and response speed.
Smart Images

Figure CN121509135B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home and smart car interconnection technology, and in particular to a device control method, apparatus, electronic device and storage medium. Background Technology
[0002] With the rapid development of IoT technology, smart homes and smart cars have gradually moved from being independent systems to interconnected systems, becoming two important components of modern smart living. Users are placing higher demands on system response speed, personalized services, and secure data interaction in their daily travel and home scenarios, especially with the growing need for cross-environment collaborative control such as "car-to-home" or "home-to-car" connections.
[0003] While several vehicle-to-home connectivity solutions have emerged in the current market, they lack effective security encryption mechanisms during data transmission, making data vulnerable to attacks and leading to user privacy leaks. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this application provides a device control method, apparatus, electronic device and storage medium.
[0005] In a first aspect, this application provides a device control method, comprising:
[0006] Receive multimodal input information from the first linkage space, and determine the target device in the second linkage space that the user intends to control based on the multimodal input information;
[0007] Obtain the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category;
[0008] Based on the multimodal input information and the device verification method, verify whether the user has control authority over the target device;
[0009] If the user has control over the target device, the target device is controlled based on the multimodal input information.
[0010] Optionally, obtaining the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category includes:
[0011] Obtain the core data type of the device associated with the target device;
[0012] In the preset correspondence between device core data types and privacy categories, determine the device privacy category corresponding to the device core data type;
[0013] In the preset correspondence between privacy categories and verification methods, determine the device verification method corresponding to the device privacy category.
[0014] Optionally, from the preset correspondence between privacy categories and verification methods, the device verification method corresponding to the device privacy category is determined, including:
[0015] If the device privacy category is high privacy category, the device verification method is a three-modal verification method;
[0016] If the device privacy category is medium privacy category, the device verification method is bimodal verification method;
[0017] If the device privacy category is low privacy category, the device verification method is single-modal verification method.
[0018] Optionally, verifying whether the user has control authority over the target device based on the multimodal input information and the device verification method includes:
[0019] If the target device's privacy category is high privacy, the user's face image, voice information, and location information are determined based on the multimodal input information. Trimodal verification is performed based on the face image, voice information, and location information. If the verification passes, it is determined that the user has control rights over the target device.
[0020] If the target device's privacy category is medium privacy category, the user's voice information and location information are determined based on the multimodal input information. Bimodal verification is performed based on the voice information and location information. If the verification is successful, it is determined that the user has control rights over the target device.
[0021] If the target device's privacy category is low privacy, the user's location information is determined based on the multimodal input information, and a single-modal verification is performed based on the location information. If the verification passes, it is determined that the user has control rights over the target device.
[0022] Optionally, controlling the target device based on the multimodal input information includes:
[0023] The multimodal input information is subjected to multimodal fusion processing to obtain multimodal fused information;
[0024] Obtain a context summary of the multimodal input information;
[0025] The device control intent is determined based on the multimodal fusion information and the context summary;
[0026] Control the target device according to the stated device control intent.
[0027] Optionally, the multimodal input information is subjected to multimodal fusion processing to obtain multimodal fused information, including:
[0028] The quality of each modal input information in the multimodal input information is evaluated using an energy function to obtain the energy uncertainty evaluation results for each mode;
[0029] Based on the energy uncertainty assessment results of each mode, the multimodal input information is dynamically fused to obtain the multimodal fused information.
[0030] Optionally, the quality of each modal input information in the multimodal input information is evaluated using an energy function to obtain the energy uncertainty evaluation results for each mode, including:
[0031] For each modal input information, the modal input information, the preset modal mean information, and the preset reconstruction error energy function are used to determine the first uncertainty assessment result;
[0032] The second uncertainty assessment result is determined by using the feature representation of the modal input information, the preset positive sample features, the preset negative sample features, and the preset contrastive learning energy function.
[0033] The modal input information is input into a preset probability model energy function to determine the third uncertainty assessment result;
[0034] The energy uncertainty assessment result is determined based on the first uncertainty assessment result, the second uncertainty assessment result, and the third uncertainty assessment result.
[0035] Optionally, dynamic multimodal fusion is performed on the multimodal input information based on the energy uncertainty assessment results of each mode to obtain the multimodal fusion information, including:
[0036] Obtain dynamic fusion weights for multiple modalities;
[0037] If the energy uncertainty assessment result of any mode is lower than a preset threshold, the historical preference weight of that mode is obtained, the dynamic fusion weight of that mode is replaced by the historical preference weight, and the dynamic fusion weight of other modes is adjusted accordingly.
[0038] The multimodal input information is fused based on the historical preference weights and the dynamic fusion weights of other modalities to obtain the multimodal fusion information.
[0039] Optionally, obtain the dynamic fusion weights of multiple modalities, including:
[0040] Feature extraction is performed on the multimodal input information to obtain modal features of multiple modalities;
[0041] Obtain the historical context of multiple modalities;
[0042] Modal features and historical context of multiple modalities are input into a preset gating function to obtain modal-level weights of multiple modalities;
[0043] Based on a preset resource-aware loss function, the modality-level weights of multiple modalities are dynamically adjusted to obtain dynamic fusion weights for multiple modalities.
[0044] Optionally, the modal features and historical context of multiple modalities are input into a preset gating function to obtain modal-level weights for multiple modalities, including:
[0045] Modal features and historical context from multiple modalities are concatenated to obtain a joint vector;
[0046] Obtain the degree of influence of the joint vector on the dynamic weights of each mode and the activation threshold of the gating function;
[0047] The joint vector, the degree of influence, and the activation threshold are input into a gating network so that the gating network maps the joint vector into dynamic weights for each modality.
[0048] The dynamic weights of each mode are normalized to obtain the modal-level weights of multiple modes.
[0049] Optionally, based on a preset resource-aware loss function, the modality-level weights of multiple modalities are dynamically adjusted to obtain dynamic fusion weights for multiple modalities, including:
[0050] The task loss value is determined based on the modal-level weights of multiple modalities;
[0051] The instruction complexity is determined based on the multimodal input information;
[0052] Obtain the complexity balance coefficient and resource complexity loss value corresponding to the instruction complexity;
[0053] The resource-aware loss value is determined based on the task loss value, complexity balance coefficient, and resource complexity loss value.
[0054] Based on the resource perception loss value, the modal-level weights of multiple modalities are dynamically adjusted to obtain the dynamic fusion weights of multiple modalities.
[0055] Optionally, determining the device control intent based on the multimodal fusion information and the context summary includes:
[0056] Obtain the context summary corresponding to the multimodal fusion information;
[0057] The multimodal fusion information and the context summary are input into a preset intent prediction model, so that the intent prediction model outputs the device control intent corresponding to the multimodal fusion information and the context summary based on a multidimensional reward strategy and a dynamic KL divergence adjustment strategy.
[0058] Secondly, this application provides a device control apparatus, comprising:
[0059] The receiving module is used to receive multimodal input information from the first linkage space and determine the target device to be controlled by the user in the second linkage space based on the multimodal input information.
[0060] The acquisition module is used to acquire the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category;
[0061] The verification module is used to verify whether the user has control authority over the target device based on the multimodal input information and the device verification method.
[0062] The control module is used to control the target device based on the multimodal input information if the user has control authority over the target device.
[0063] Thirdly, this application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0064] Memory, used to store computer programs;
[0065] The processor, when executing a program stored in memory, implements the device control method described in any of the first aspects.
[0066] Fourthly, this application provides a computer-readable storage medium storing a program for a device control method, wherein when the program for the device control method is executed by a processor, it implements the steps of any of the device control methods described in the first aspect.
[0067] The beneficial effects of this invention are as follows: This application's embodiments determine the target device based on multimodal input information and verify the user's control permissions according to the target device's device privacy category and device verification method. When the user has control permissions over the target device, the target device is controlled based on the multimodal input information. The user's control permissions are then hierarchically verified according to the target device's device privacy category, ensuring that the user's control permissions match the device privacy category of the target device to be controlled. This avoids the leakage of privacy content of the target device that does not match the user's control permissions, thereby improving the security and privacy protection level of vehicle-home interconnection. Attached Figure Description
[0068] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0069] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 An architecture diagram of a vehicle-home interconnection system provided in this application embodiment;
[0071] Figure 2 A flowchart of a device control method provided in an embodiment of this application;
[0072] Figure 3 for Figure 2 An exemplary flowchart of step S103;
[0073] Figure 4 for Figure 2 Flowchart of step S104;
[0074] Figure 5 for Figure 4 Flowchart of step S201;
[0075] Figure 6 for Figure 5 Flowchart of step S302;
[0076] Figure 7 for Figure 6 Flowchart of step S401;
[0077] Figure 8 A structural diagram of a device control apparatus provided in an embodiment of this application;
[0078] Figure 9 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0080] Because existing vehicle-to-home (V2X) solutions lack effective security encryption mechanisms during data transmission, data is vulnerable to attacks, leading to user privacy leaks. Therefore, this application provides a device control method, apparatus, electronic device, and storage medium. This enables deep intelligent collaboration and personalized services between the vehicle and the home environment, reducing the number of user-initiated steps. Furthermore, the system introduces a security encryption mechanism to ensure the security and privacy protection of data transmission during V2X; simultaneously, it optimizes the data processing path through an edge computing architecture, improving data processing efficiency and response speed in V2X scenarios.
[0081] This application relates to the intersection of vehicle networking and smart home technology, specifically to a vehicle-home interconnection system integrating dynamic encrypted communication across the vehicle, cloud, and home terminals, multimodal scene recognition, smart device command distribution, and AI voice interaction. It is suitable for scenarios involving the coordinated control of new energy vehicles and smart home devices, safety alerts, and disaster warnings. Figure 1 As shown, the system mainly involves three entities: the vehicle-side T-Box (integrating a voiceprint module, a 5G communication module, and a multimodal interactive sensor), the cloud platform (an IoT management platform including a permission engine and a large scene reasoning model), and the home-side (smart gateway connection and smart home terminal).
[0082] The T-Box communicates directly with the cloud platform via a 5G network. The cloud platform then connects to a smart gateway at home via the internet, forming a communication link between the vehicle, the cloud, and the home. In terms of software architecture, the system primarily consists of two core processing modules: a security encryption module and a dynamic intent recognition module. The former ensures the integrity and confidentiality of data during transmission and interaction, while the latter uses multimodal information fusion and inference algorithms to achieve real-time recognition and intelligent response to user intent.
[0083] like Figure 1As shown, the overall system data flow has a two-way interactive structure, and the specific process is as follows: First, the vehicle-side T-Box collects location, voice, and other data in real time and uploads it to the cloud platform via the 5G network. After completing data parsing and dynamic intent recognition, it generates corresponding control commands and sends them to the home gateway, while providing security verification and intelligent decision support. Then, the home terminal uploads voice data, image data, etc., to the cloud platform, which also analyzes and processes them, and sends corresponding control commands to the vehicle terminal to complete remote control of the vehicle. Throughout the process, the cloud platform always acts as the data hub and decision-making core, responsible for information aggregation, security encryption verification, and collaborative control strategy generation to ensure the stable, intelligent, and secure operation of the vehicle-home system. Specifically, the vehicle maintains communication with the cloud through the T-Box. The cloud receives vehicle status information from the vehicle terminal, which is used for security encryption and as the driving condition for vehicle-home dynamic intent reasoning. On the other hand, the cloud also maintains communication with the vehicle control domain through the vehicle's T-Box using the MCU, receiving vehicle control signals sent from the cloud to drive the vehicle control domain and execute vehicle control commands.
[0084] The vehicle terminal collects vehicle information, including one or more of vehicle location information, voice information, and image information. It is responsible for triggering the vehicle-to-home control function according to preset trigger conditions, such as the voice command "I want to control the air conditioner at home", and uploading the corresponding information to the cloud to prepare for subsequent data processing and judgment in the cloud.
[0085] The aforementioned body control domain is the vehicle body control terminal, responsible for the functional control of components such as doors, windows, seats, and air conditioning. It is used to receive or detect signals such as switches and to perform communication and adjustment functions.
[0086] The cloud includes vehicle cloud and IoT cloud, which accept various user device information and control commands uploaded by vehicle terminals or home gateways, and perform data processing and command issuance according to preset security encryption logic and dynamic intent reasoning logic.
[0087] Based on the aforementioned system, embodiments of this application provide a device control method, such as... Figure 2 As shown, it includes:
[0088] Step S101: Receive multimodal input information from the first linkage space, and determine the target device in the second linkage space that the user intends to control based on the multimodal input information;
[0089] In this embodiment, the first linkage space can be a vehicle space, a home space, or an office space, etc. The second linkage space can be a different space from the first linkage space. When the first linkage space is a vehicle space, the second linkage space can be a home space or an office space, etc., and when the first linkage space is a home space or an office space, the second linkage space can be a vehicle space.
[0090] Multimodal input information is collected by multiple sensors (such as microphones, cameras, and GPS positioning modules) within the first linkage space. Multimodal input information includes one or more of the following: sound information, video / image information, location information, gesture information, and device status information. For example, a user can input the target device to be controlled through voice information. At the same time as voice information is collected, video / image information, time information, and location information are also collected as multimodal input information.
[0091] The system receives one or more multimodal signals (such as voice commands, image data, gesture signals, sensor data, etc.). For example, if a user says, "After picking up my child from school tonight, turn on the air conditioner in advance," the system will simultaneously collect one or more data, such as voice information, time information (Friday afternoon), and device status information (air conditioner location and status).
[0092] The environmental perception initialization preprocesses data through the environmental perception module, which can identify the scene type (such as vehicle-home interconnection, home control) and the complexity of the command (simple commands such as "turn on the air conditioner", complex commands such as "turn on the air conditioner in advance after picking up the child from school tonight").
[0093] Assume the system receives three modal inputs: speech modality (m1): acoustic feature vectors of user voice commands, such as MFCC or Mel spectrum, with dimension d1; historical preference modality (m2): statistical features of user's historical temperature preferences (such as mean and variance), with dimension d2; gesture modality (m3): keypoint coordinates or classification labels for gesture recognition, with dimension d3. The input data can be represented as follows: ,in .
[0094] Step S102: Obtain the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category;
[0095] In this embodiment, device privacy categories include: high privacy category, medium privacy category, and low privacy category. Device privacy categories that can be linked and controlled in the first linkage space and the second linkage space can be preset, and device verification methods can be preset for each device privacy category. As shown in Table 1, the high privacy category mainly includes terminal devices that collect, store, or transmit highly sensitive information such as facial recognition information, fingerprint data, and home space images; the medium privacy category mainly refers to devices involving user location trajectory and environmental temperature and humidity data; and the low privacy category mainly includes control devices such as entertainment systems and lighting control systems.
[0096] Table 1: Equipment Categories and Relevant Standards
[0097]
[0098] Based on the aforementioned privacy categories, a tiered and dynamic verification strategy is implemented by combining biometric information with geographic location information. The verification logic is as follows: devices in the high privacy category (such as vehicle terminals) employ full trimodal verification using face, voiceprint, and location information; devices in the medium privacy category (such as home gateways) employ bimodal verification using voiceprint and location information; and devices in the low privacy category (such as environmental sensors) perform single verification based solely on location information, thus improving the success rate of defending against GPS spoofing attacks.
[0099] This verification engine adopts a hierarchical decision-making architecture, which can intelligently combine and call the corresponding verification methods according to the privacy level of the device, and realize adaptive adjustment of verification strength, thereby optimizing user experience and resource allocation while ensuring the overall security of the system.
[0100] In one embodiment of this application, step S102, obtaining the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category, includes:
[0101] Obtain the core data type of the target device; determine the device privacy category corresponding to the core data type in the preset correspondence between core data types and privacy categories (as shown in Table 1); determine the device verification method corresponding to the privacy category in the preset correspondence between privacy categories and verification methods.
[0102] Specifically, in the preset correspondence between privacy categories and verification methods, the device verification method corresponding to the device privacy category is determined, including: if the device privacy category is a high privacy category, the device verification method is a trimodal verification method; if the device privacy category is a medium privacy category, the device verification method is a bimodal verification method; if the device privacy category is a low privacy category, the device verification method is a unimodal verification method.
[0103] Step S103: Verify whether the user has control authority over the target device based on the multimodal input information and the device verification method;
[0104] In one embodiment of this application, such as Figure 3 As shown, step S103 verifies whether the user has control authority over the target device based on the multimodal input information and the device verification method, including:
[0105] If the target device's privacy category is high privacy, the user's facial image, voice information, and location information are determined based on the multimodal input information. Trimodal verification is then performed based on the facial image, voice information, and location information. If the verification passes, the user is determined to have control over the target device. If the target device's privacy category is medium privacy, the user's voice information and location information are determined based on the multimodal input information. Bimodal verification is then performed based on the voice information and location information. If the verification passes, the user is determined to have control over the target device. If the target device's privacy category is low privacy, the user's location information is determined based on the multimodal input information. Unimodal verification is then performed based on the location information. If the verification passes, the user is determined to have control over the target device. If facial verification, voiceprint verification, and location verification fail, a warning message is sent to the registered account.
[0106] The face recognition logic flow is as follows: First, a pre-trained model (such as ResNet) is loaded using OpenCV's DNN module; then, face feature vectors (128-dimensional floating-point arrays) are extracted; finally, the cosine similarity between the face features and those of the registered template is calculated, with the core formula as follows:
[0107]
[0108] Where Similarity represents cosine similarity, A is the feature vector to be compared, and B is the facial feature vector of the registered template. The dot product of vectors reflects the consistency of direction. and Let be the vector magnitude.
[0109] This formula calculates the cosine of the angle between two vectors, used to measure the similarity of feature vectors in face recognition. Its range is ( ). The higher the value (1, 1), the higher the similarity. Based on the above calculation results, a threshold is determined. When the similarity result is greater than or equal to 85%, the face verification passes.
[0110] The voiceprint processing part of the aforementioned voiceprint recognition takes into account the limitations of the in-vehicle environment and the in-vehicle microphone itself, and needs to reduce the impact of noise in the in-vehicle environment. Therefore, the preprocessing steps include noise detection, Wiener filtering and low-pass filtering. At the same time, it considers combining Mel frequency cepstral coefficient (MFCC) feature extraction and cosine similarity for result matching to improve the recognition accuracy in the noisy environment of the vehicle cabin.
[0111] The voiceprint recognition module was designed with full consideration of the complex noise levels and limited microphone performance in the vehicle environment. To reduce the impact of environmental noise on recognition accuracy, this application introduces noise detection, Wiener filtering, and low-pass filtering in the preprocessing stage of voiceprint recognition to perform multi-level noise reduction on the original speech signal and improve its quality. In the feature extraction stage, the system uses the MFCC algorithm to extract voiceprint feature parameters. Finally, the extracted features are matched with the user's voiceprint template using a cosine similarity algorithm to complete identity verification.
[0112] To address the noise detection problem in vehicle-mounted voiceprint recognition, considering that the traditional energy thresholding method is simple to implement but lacks robustness, while spectral analysis, although computationally complex, has a stronger ability to distinguish noise from speech, a hybrid noise detection method based on energy thresholding and spectral analysis is proposed for the noise detection section, balancing detection efficiency and accuracy. Specifically, the method includes the following steps: First, the input continuous audio signal is segmented into multiple short frames for frequency domain analysis; second, the frame signals are initially screened using a dynamic energy thresholding method for energy threshold detection; third, spectral feature analysis is performed to distinguish noise from speech based on spectral complexity; finally, a comprehensive judgment is made based on the energy and spectral feature information for each frame.
[0113] The dynamic energy threshold detection involves adaptive adjustment of the permission threshold. The specific logic of the dynamic threshold adjustment strategy in dynamic energy threshold detection is as follows: First, an initial energy threshold is calculated based on the data from the previous 3 seconds of silence. Then, the threshold is adaptively updated every 5 seconds based on the audio data from the most recent second. Finally, smoothing is performed using a moving average filter (window length = 5 frames), expressed by the following formula:
[0114]
[0115] in, For the updated threshold, For historical thresholds, This is the smoothing coefficient (typical value 0.95). For the first To calculate the number of frames within the window.
[0116] Based on the aforementioned vehicle environment, an optimized MFCC feature extraction method is proposed, taking into account the vehicle noise environment and further subdividing parameters such as frame length, frame shift, and number of filter banks.
[0117] The specific optimization extraction process for MFCC feature extraction is as follows: First, input audio data as the original signal; then, perform pre-emphasis filtering on the audio signal. This step uses the librosa.effects.preemphasis function with a coefficient set to 0.95.
[0118] Further frame segmentation and windowing were performed using Hamming windows, with a frame length of 400 (corresponding to...). The frame shift is 160 (10ms). Framing is used to process continuous signals into short-term stable frames, and windowing (such as Hamming window) is used to reduce spectral leakage.
[0119] Furthermore, power spectrum calculation is performed, and the time-domain signal is converted into frequency-domain energy distribution through short-time Fourier transform (STFT); furthermore, a Mel filter bank is used to simulate the human ear's perception of different frequencies, converting linear frequencies into Mel frequencies;
[0120] Further logarithmic energy calculation is used to take the logarithm of the energy after Mel filtering and add a minimum value to prevent zero from occurring when taking the logarithm, thereby compressing the dynamic range and simulating the nonlinear perception of sound intensity by the human ear.
[0121] Further, the Discrete Cosine Transform (DCT) is used to retain the first 20 coefficients, which are the MFCC eigenvectors.
[0122] Further calculations of first-order and second-order differences (delta and delta-delta) capture the dynamic features of speech.
[0123] When the real-time extracted MFCC feature vector (including differential features) is compared with the template features in the pre-stored voiceprint feature library, if the result exceeds the preset threshold, the voiceprint verification can be determined to be successful.
[0124] For example, the parameter adjustments for parameters such as frame length, frame shift, and number of filter banks are shown in Table 2:
[0125] Table 2: MCFF Model Parameter Table
[0126]
[0127] The location verification module's specific logic includes: First, obtaining the GPS coordinates (latitude and longitude) of the vehicle and the terminal device corresponding to the registered account; then, after correcting for Earth's curvature using the Haversine formula, calculating the distance between the vehicle's location and the terminal device's location. If the distance between the two is less than or equal to 5 meters, the verification is considered successful. The Haversine formula used to calculate the distance between the vehicle's location and the terminal device's location is as follows:
[0128]
[0129]
[0130]
[0131] in, As an intermediate variable, For two points, latitude (radians) Due to latitude difference, Due to longitude difference, Angular distance R represents the distance between the vehicle's location and the terminal equipment's location, where R is the Earth's radius (6371 kilometers). The great circle distance between the two points.
[0132] Step S104: If the user has control authority over the target device, control the target device based on the multimodal input information.
[0133] This application embodiment determines the target device based on multimodal input information and verifies the user's control permissions according to the target device's device privacy category and device verification method. When the user has control permissions over the target device, the target device is controlled based on the multimodal input information. The user's control permissions are then hierarchically verified according to the target device's device privacy category to ensure that the user's control permissions match the device privacy category of the target device to be controlled. This avoids the leakage of privacy content of the target device that does not match the user's control permissions, thereby improving the security and privacy protection level of vehicle-home interconnection.
[0134] Since vehicle-to-home connectivity primarily relies on fixed communication protocols and preset instruction sets, it is difficult to achieve dynamic, user-intent-based intelligent collaboration. Furthermore, traditional edge computing methods suffer from low computational efficiency and slow response speed when processing complex data in vehicle-to-home connectivity scenarios. Therefore, in another embodiment of this application, as... Figure 4 As shown, step S104 controls the target device based on the multimodal input information, including:
[0135] Step S201: Perform multimodal fusion processing on the multimodal input information to obtain multimodal fused information;
[0136] In this embodiment, the multimodal fusion processing includes: dynamic multimodal fusion (DynMM) and quality-aware multimodal fusion (QMF). The dynamic multimodal fusion uses a gating function and a resource-aware loss function to adaptively fuse multimodal information based on the input data, reducing computational costs (up to 46.5%) while maintaining accuracy.
[0137] The specific gating decision-making mechanism introduces learnable gating functions to make dynamic decisions at the modality or fusion level for multimodal inputs such as voice, images, and gestures. For example, in the vehicle-home interconnection scenario, when a user adjusts the in-vehicle temperature via voice command, the system prioritizes fusing voice modality with historical temperature preference data and temporarily suspends processing irrelevant gesture signals.
[0138] The specific resource-aware loss function encourages the model to choose more computationally efficient modal combinations by penalizing complex fusion paths. For example, single-modal inference is used for simple instructions (such as "turn on the air conditioner"), while multimodal deep fusion is enabled for complex instructions (such as "adjust the route home according to my schedule").
[0139] The aforementioned Quality-Aware Multimodal Fusion (QMF) utilizes energy-based uncertainty to assess modal quality, dynamically adjusts the fusion strategy, and enhances generalization ability. Specifically, it evaluates the data quality of each modality based on an energy function and dynamically adjusts the weights. For example, in noisy environments, the weight of microphone signals is reduced, increasing the decision-making proportion of the visual modality.
[0140] The specific process of the aforementioned context modeling technique is as follows: First, a mandatory structured response format is adopted to guide the model to explicitly summarize multimodal context and perform reflective logical reasoning, effectively avoiding shortcut problems; second, based on a large model-driven multidimensional reward mechanism, context rewards, format rewards, and accuracy rewards are used to guide the model to improve its context understanding ability and integrate advanced reasoning strategies such as reflection, deduction, and induction; finally, a dynamic KL divergence adjustment strategy is introduced to dynamically balance the exploratory nature and stability of the model during training, avoiding excessive model constraint or divergence.
[0141] The specific context summary module and reverse logic reasoning are used, where the context summary template requires the model to explicitly generate a context summary before reasoning. For example, if a user says "I want to watch a movie when I get home tonight," the system needs to first summarize historical behavior (the user often uses the home theater on weekend evenings) and environmental state (the current time is Friday afternoon).
[0142] Reflective logical reasoning combines LLM to assess whether the reasoning process incorporates higher-level logic such as reflection and deduction. For example, when determining user intent, it is necessary to verify whether it is related to multi-dimensional information such as schedules and device status.
[0143] The LLM-driven multidimensional reward mechanism includes context rewards and logical rewards. Context rewards are given to models that accurately cover key information based on the matching degree between the generated context summary and the real scene. For example, in vehicle-home connectivity, a high score is given if the model correctly identifies the user's schedule association of "picking up the child from school and going home." Logical rewards assess whether the reasoning process conforms to causal relationships. For example, the user command "lower the car's interior temperature" needs to be associated with historical preferences (the user usually sets it to 26℃) and environmental data (the current outside temperature is 30℃).
[0144] The specific dynamic KL divergence training strategy primarily balances exploratory and stability. In the early stages of training, it allows the model to freely explore multimodal combinations, while later, KL divergence constraints prevent divergence. For example, when deploying on edge devices, multiple fusion strategies can be tried initially, and then a fixed, efficient path can be established later.
[0145] In one embodiment of this application, step S201 performs multimodal fusion processing on the multimodal input information to obtain multimodal fused information, such as... Figure 5 As shown, it includes:
[0146] Step S301: Use the energy function to evaluate the quality of each modal input information in the multimodal input information, and obtain the energy uncertainty evaluation result of each mode;
[0147] Methods for calculating energy uncertainty include: an energy function based on reconstruction error: assuming the modal data should satisfy a certain latent distribution (such as a Gaussian distribution), energy is defined by calculating the difference between the input data and the reconstructed data (or the mean). This method is suitable for data with a clear latent distribution, such as speech signals (assuming a Gaussian noise model) and image pixels (assuming a locally smooth distribution).
[0148] Energy function based on contrastive learning: Uncertainty is defined by comparing the energy difference between positive samples (reliable data) and negative samples (unreliable data). This calculation method is suitable for data with semantic or categorical information: such as sentiment analysis (happy / angry speech), gesture recognition (clenched fist / waving); it is also suitable for distinguishing between positive and negative samples: contrastive learning distinguishes reliable data (positive samples) from unreliable data (negative samples); and it is suitable for modal data with high-dimensional features: such as speech / image features extracted by deep learning, where contrastive learning can capture the similarity between features. Typical task: multimodal sentiment analysis (multimodal fusion of speech, text, and facial expressions).
[0149] Energy function based on probabilistic models: Modal data is modeled as a probability distribution (such as a Gaussian mixture model), and energy is defined by calculating the log-likelihood of the data. This method is suitable for data with complex distributions, such as multimodal distributions (gesture recognition may contain multiple gesture categories); data requiring uncertainty quantification: probabilistic models (such as Gaussian mixture models) can provide the probability of data belonging to each component, directly reflecting uncertainty; and data with ambiguity, such as speech recognition (accents or dialects cause the probability distribution to be dispersed).
[0150] In one embodiment of this application, step S301 uses an energy function to evaluate the quality of each modal input information in the multimodal input information, and obtains the energy uncertainty evaluation result of each modality, including:
[0151] For each modal input information, a first uncertainty assessment result is determined by combining the modal input information, preset modal mean information, and preset reconstruction error energy function; a second uncertainty assessment result is determined by combining the feature representation of the modal input information, preset positive sample features, preset negative sample features, and preset contrastive learning energy function; a third uncertainty assessment result is determined by inputting the modal input information into a preset probability model energy function; and an energy uncertainty assessment result is determined based on the first, second, and third uncertainty assessment results.
[0152] The formula for the reconstruction error energy function is as follows:
[0153]
[0154] in, For input data, This is the mean (or reconstructed value) of the data.
[0155] Example: Speech modality: If the speech signal Compared to the mean of clean speech If the difference is large (e.g., background noise is present), then High indicates high uncertainty; if the difference is small (such as pure speech), then... Low indicates low uncertainty; Image modality: If the IoU (Intersection over Union) of the detected bounding box differs greatly from the ground truth bounding box, the reconstruction error is high. High indicates high detection uncertainty.
[0156] The formula for the contrastive learning energy function is as follows:
[0157]
[0158] in, For input data, Feature representation of input data For negative samples, As a positive sample, T The temperature parameter is used to scale the similarity scores of sample pairs. This represents the inner product similarity between the current sample and the positive sample (the larger the value, the more similar the sample).
[0159] Example: In multimodal sentiment analysis, if the features of the speech modality... Positive sample features of happy expressions High similarity Low indicates low uncertainty in emotional judgment; if compared with negative sample features of angry expressions High similarity "High" indicates high uncertainty.
[0160] The formula for the energy function of the probabilistic model is as follows:
[0161]
[0162] in, For input data, For the first The mixing coefficients (weights) of a Gaussian distribution. For the first The mean vector of a Gaussian distribution (cluster centers). For the first The covariance matrix (distribution shape) of a Gaussian distribution. No. The probability density function of a Gaussian distribution.
[0163] Example: In gesture recognition, if the log-likelihood of gesture data... High (i.e., the data matches the training distribution), then A low log-likelihood indicates low uncertainty in identification; if the log-likelihood is low (e.g., the data deviates from the training distribution), then... "High" indicates high uncertainty.
[0164] Since the above three methods quantify energy uncertainty from different perspectives (reconstruction error reflects noise, contrastive learning reflects semantic differences, and probabilistic models reflect distribution ambiguity), their combined use can improve the comprehensiveness of the assessment. A weighted fusion strategy can be adopted as follows to sum the three energy values using a weighted approach to obtain the comprehensive energy.
[0165]
[0166] in, The hyperparameters are the weights of each energy function. The determination will be adaptively adjusted: based on online learning in a dynamic environment, the specific determination steps are as follows:
[0167] 1. Initialize weights: Set initial weights using manual parameter tuning or data-driven methods.
[0168] 2. Online rule update: First, use exponential moving average (EMA) to dynamically adjust the weights based on the performance of new data; then further reinforcement learning treats the weights as actions and learns the optimal policy through reward functions (such as improving evaluation metrics).
[0169] 3. Sliding window validation: Only the data from the most recent N samples are used to update the weights to avoid interference from historical data.
[0170] Step S302: Based on the energy uncertainty assessment results of each mode, perform dynamic multimodal fusion on the multimodal input information to obtain the multimodal fusion information.
[0171] The core objective of this phase is to integrate dynamic data from different modalities, such as vision, text, voice, and sensors, into a unified feature representation, thereby solving the "data silo" problem.
[0172] In the dynamic multimodal fusion architecture, the gating decision mechanism, resource-aware loss function, and quality-aware fusion (QMF) work together through three dimensions of dynamism, efficiency, and robustness to form a complete adaptive fusion framework.
[0173] In one embodiment of this application, step S302 involves dynamic multimodal fusion of the multimodal input information based on the energy uncertainty assessment results of each modality, to obtain the multimodal fused information, such as... Figure 6 As shown, it includes:
[0174] Step S401: Obtain the dynamic fusion weights of multiple modalities;
[0175] In one embodiment of this application, step S401 obtains the dynamic fusion weights of multiple modalities, such as... Figure 7 As shown, it includes:
[0176] Step S501: Extract features from the multimodal input information to obtain modal features of multiple modalities;
[0177] For example, the modal features of multiple modalities can be speech MFCC features, image CNN features, point cloud voxelization features, etc.
[0178] Step S502: Obtain the historical context of multiple modalities;
[0179] Step S503: Input the modal features and historical context of multiple modalities into a preset gating function to obtain the modal-level weights of multiple modalities;
[0180] In one embodiment of this application, step S503 inputs the modal features and historical context of multiple modalities into a preset gating function to obtain modal-level weights of multiple modalities, including: concatenating the modal features and historical context of multiple modalities to obtain a joint vector; obtaining the degree of influence of the joint vector on the dynamic weights of each modality and the activation threshold of the gating function; inputting the joint vector, the degree of influence, and the activation threshold into a gating network so that the gating network maps the joint vector to the dynamic weights of each modality; and normalizing the dynamic weights of each modality to obtain modal-level weights of multiple modalities.
[0181] In practical applications, gating networks generate modal weights through the following steps. The specific steps are as follows:
[0182] 1. Feature concatenation: Concatenate all modal features into a joint vector. ;
[0183] 2. Gated Networks: These networks map joint vectors to modal dynamic weights through fully connected layers. The specific formula is as follows:
[0184]
[0185] in, For modality Dynamic weights (such as voice, historical preferences, gestures, etc.). The sigmoid function simulates the nonlinear decision-making of a "gated switch," automatically deciding whether to retain a mode based on modal quality, task requirements, or contextual information. The output range is... 0 indicates completely off, 1 indicates fully on. Control input features For modes The degree of influence of the weights; The input feature vector contains current modality information and historical context (such as user temperature preferences and environmental conditions). As a bias term, adjust the activation threshold of the gating function;
[0186] 3. Weight Normalization: Complex instructions (such as "Adjust the route according to the schedule, pick up the child from school tonight and go home, turn on the air conditioner in advance") require fusion-level decision-making, which can be achieved through Softmax to ensure... For simple instructions ("turn on the air conditioner") that are modal-level decisions (such as masking irrelevant modes), the Sigmoid output is used directly.
[0187] The core task of the gating mechanism is to dynamically select modes, and its output is the weights of each mode. This process generates modality-level or fusion-level dynamic decisions based on real-time features of the input data (such as speech clarity and gesture confidence) and task requirements (such as the urgency of temperature regulation) through learnable functions. For example, in a temperature regulation task, the system might output... =0.8 (voice) =0.7 (historical preference) =0.1 (gesture), indicating that the gesture signal was significantly suppressed.
[0188] Step S504: Based on the preset resource-aware loss function, the modal-level weights of multiple modalities are dynamically adjusted to obtain the dynamic fusion weights of multiple modalities.
[0189] In one embodiment of this application, step S504 dynamically adjusts the modal-level weights of multiple modalities based on a preset resource-aware loss function to obtain dynamic fusion weights for multiple modalities, including:
[0190] The task loss value is determined based on the modal-level weights of multiple modalities; the instruction complexity is determined based on the multimodal input information; and the task loss value is obtained.
[0191] The complexity balance coefficient and resource complexity loss value corresponding to the instruction complexity are determined; the resource awareness loss value is determined based on the task loss value, complexity balance coefficient and resource complexity loss value; the modality-level weights of multiple modalities are dynamically adjusted according to the resource awareness loss value to obtain the dynamic fusion weights of multiple modalities.
[0192] The resource-aware loss function encourages the model to choose computationally more efficient modal combinations by penalizing complex fusion paths. This achieves a balance between efficiency and accuracy in resource-constrained devices. For example, single-modal inference is used for simple instructions (such as "turn on the air conditioner"), while multimodal deep fusion is enabled for complex instructions (such as "turn on the air conditioner before picking up the child from school tonight"). The specific formula is as follows:
[0193]
[0194] in, For task-related losses (such as the MSE loss of temperature regulation), the result is a fusion of modal-level weights from multiple modes of gating decision. This is a complexity balancing coefficient. A value of 0 indicates that only task loss is optimized, ignoring resource consumption (encouraging complex reasoning); a value greater than 0 introduces resource penalties, and the larger the value, the more emphasis is placed on computational efficiency (encouraging lightweight paths). To compensate for resource complexity losses (such as computational latency and energy consumption regularization terms), the gating weights are directly constrained. .
[0195] The goal of resource-aware loss is to optimize system resource allocation. Its inputs are the output of gating decisions (e.g., modal weights) and the state of system resources (e.g., computational latency, energy consumption). Through loss function constraints, the model minimizes resource consumption (e.g., reducing the computational overhead of low-weight modalities) while ensuring task performance. For example, if the gesture modality is assigned low weight by the gating mechanism (…), =0.1), the resource-aware loss will further penalize its computational resource consumption, encouraging the model to skip gesture feature extraction in subsequent inference.
[0196] During training, the gating mechanism... Learning the optimal mode combination, and Then the distribution of gating weights is dynamically adjusted. For example, if the gesture modality contributes little to the task ( High), and its computational cost is large ( (High), then This will cause the model to decrease Conversely, if the gesture modality is critical to the task (such as an emergency stop gesture), even if resource consumption is high, It will also allow the retention of its weight.
[0197] Step S402: If the energy uncertainty assessment result of any mode is lower than a preset threshold, obtain the historical preference weight of that mode, replace the dynamic fusion weight of that mode with the historical preference weight, and adjust the dynamic fusion weight of other modes accordingly.
[0198] If the historical preference weight is low, the dynamic fusion weights of other modalities are increased. For example, in a noisy environment, the confidence score of the speech modality decreases, and QMF will... The weight of historical preferences has been reduced from 0.8 to 0.4, while the weight of visual modalities (such as dashboard readings) has been increased.
[0199] Step S403: Based on the historical preference weights and the dynamic fusion weights of other modalities, the multimodal input information is fused to obtain the multimodal fusion information.
[0200] Based on the aforementioned gating decision-making mechanism, resource-aware loss function, and quality-aware fusion (QMF) synergy, a dynamic, efficient, and robust triangular closed loop is formed. Specifically:
[0201] 1. Coordination between gating decision-making and resource awareness
[0202] Forward association: Modal weights for gating decision generation The resource-aware loss function is directly input as the basis for calculating cost allocation.
[0203] Backward optimization: Resource-aware loss adjusts the gating network parameters through backpropagation, encouraging the model to choose low-cost modal combinations while ensuring task performance.
[0204] Example: If gesture modalities contribute little to the temperature regulation task and have high computational cost, resource-aware loss will reduce their weight. 手势 This causes the gating mechanism to skip gesture feature extraction in subsequent inference.
[0205] 2. Synergy between Gated Decision Making and QMF
[0206] Quality-Driven Gating Adjustment: The uncertainty assessment results of QMF can dynamically override the initial weights of the gating decision. Example: In noisy environments, the confidence score of the speech modality decreases, and QMF will... 语音 The weighting was reduced from 0.8 to 0.4, while the weighting of visual modalities (such as dashboard readings) was increased.
[0207] Modal selection for task adaptation: The gating mechanism, combined with QMF output, prioritizes modalities that are both high-quality and task-relevant. Example: In the task "After picking up the child from school tonight, turn on the air conditioner in advance," even if the voice modal quality is high, if the schedule data is obtained using historical setting preferences, the gating mechanism will still increase the weight of the historical setting preference modal.
[0208] 3. The trade-off between resource awareness and QMF in terms of efficiency and robustness: Resource awareness loss encourages the use of low-cost modalities, while QMF ensures that low-quality modalities are not incorrectly relied upon. Example: In simple instructions (such as "turn on the air conditioner"), resource awareness loss drives the model to use a single modality (speech), while QMF forces a switch to app control (visual modality) when there is a lot of speech noise.
[0209] The output of this stage is the generation of structured multimodal feature vectors, which serve as input for contextual modeling. For example, in a smart car-home interconnection scenario, user voice, text input, and historical dialogue records are integrated to form comprehensive features that include emotion, intent, and historical behavior.
[0210] Step S202: Obtain the context summary of the multimodal input information;
[0211] Context modeling builds a high-level understanding of scenarios, tasks, or user intentions based on multimodal feature tensors, thus solving the problem of "fragmented information".
[0212] The context summary template requires the model to explicitly generate a context summary before inference. For example, if a user says, "After picking up my child from school tonight, I'll turn on the air conditioner beforehand," the system needs to summarize the user's past behavior and environmental state. The specific steps are as follows:
[0213] 1. Input the user's original command: "After picking up the child from school tonight, turn on the air conditioner beforehand."
[0214] 2. Retrieve historical data from the engine: User schedule: 17:00 today, elementary school ends; Device status: Air conditioning is currently off, historical preference is 26℃; Environmental data: Current outside temperature is 30℃, indoor temperature is 28℃.
[0215] Pseudocode example of using LLM to generate structured summaries:
[0216] def generate_context_summary(user_input, history_data):
[0217] prompt = f"""
[0218] User command: {user_input}
[0219] Historical data: schedule {history_data['schedule']}, device {history_data['device_status']}
[0220] Environment status: Temperature {history_data['environment']}
[0221] Please generate a JSON digest that includes the scenario type, key equipment, and time constraints.
[0222] """
[0223] summary = llm_model(prompt)
[0224] # Output example:
[0225] # {"scene_type": "Pick up the kids from school", "target_device": "Air conditioner", "time_constraint": "After 5:00 PM", "temp_ref": 26}
[0226] return parse_json(summary)
[0227] Step S203: Determine the device control intent based on the multimodal fusion information and the context summary;
[0228] In one embodiment of this application, step S203, determining the device control intent based on the multimodal fusion information and the context summary, includes:
[0229] Obtain the context summary corresponding to the multimodal fusion information; input the multimodal fusion information and the context summary into a preset intent prediction model, so that the intent prediction model outputs the device control intent corresponding to the multimodal fusion information and the context summary based on a multidimensional reward strategy and a dynamic KL divergence adjustment strategy.
[0230] The multi-dimensional reward mechanism includes context rewards and logical rewards. Context rewards are given to models that accurately cover key information based on the matching degree between the generated context summary and the real scene. For example, in vehicle-home connectivity, a high score is given if the model correctly identifies the user's schedule association of "picking up the child from school and going home." Logical rewards assess whether the reasoning process conforms to causal relationships. For example, the user command "turn on the air conditioner in advance" needs to be associated with historical preferences (the user usually sets it to 26℃) and environmental data (the current outside temperature is 30℃).
[0231] The main function of the dynamic KL divergence training strategy is to balance exploration and stability. In the early stages of training, the model is allowed to freely explore multimodal combinations, while later, KL divergence constraints prevent divergence. For example, consider the scenario of "turning on the air conditioner before picking up the child from school tonight." Based on the previous analysis of the ideal air conditioner temperature setting, KL divergence analysis shows that the current strategy distribution shows 70% of cases where the air conditioner is turned on 20 minutes in advance, 20% 30 minutes in advance, and the reference distribution shows 80% 20 minutes in advance. Therefore, the calculated KL divergence is 0.15 < the threshold of 0.2, thus maintaining the current exploration rate of 15%.
[0232] Through this dynamic KL divergence control, the system can fully explore multimodal combination strategies (such as different time offsets + temperature settings) in the early stages of training, accumulating high-reward samples. In the later stages, KL divergence constraints ensure that the strategy remains stable near the reference distribution, preventing decision divergence due to overexploration. Ultimately, in the scenario of "picking up the child and turning on the air conditioner," the system will stably output a high-confidence strategy of "turning on the air conditioner at 26°C 20 minutes in advance," while reserving 15% of the exploration space to cope with environmental changes (such as dynamic adjustment when the outside temperature suddenly rises to 35°C).
[0233] Reflective logical reasoning combines LLM (Leadership Management) to assess whether the reasoning process incorporates higher-level logic such as reflection and deduction. For example, when determining user intent, it's necessary to verify whether it's linked to multi-dimensional information such as schedules and device status. The verification logic is as follows:
[0234] 1. Multidimensional information association verification:
[0235] Check if the schedule is linked (school dismissal time matches);
[0236] Check the status of the associated devices (air conditioner is currently off);
[0237] Check if it is related to environmental data (temperature adjustment requirement triggered by 30℃ outside the vehicle);
[0238] 2. Causal relationship verification, code example as follows:
[0239] # Verify partial timing logic of "picking up the child → going home → adjusting the temperature"
[0240] def validate_causal_chain(summary):
[0241] If "picking up the child" is in the summary and "going home" is in the summary:
[0242] If "temperature adjustment" is in the summary and "outside temperature 30℃" is in the summary:
[0243] return True # Consistent with causal logic
[0244] return False.
[0245] Based on the fused multimodal information and context summary, the final intent is output (e.g., "set the home air conditioner to 26°C at 17:20").
[0246] For ease of understanding, this application also provides a typical example of dynamic intent recommendation process (vehicle-home interconnection scenario).
[0247] 1. User voice commands:
[0248] "After picking the kids up from school tonight, turn on the air conditioner before you get home."
[0249] 2. Multimodal fusion:
[0250] Prioritize the integration of voice (intent core) and historical schedule data (time to pick up children).
[0251] Temporarily suspend processing irrelevant gesture signals (such as a user scratching their head);
[0252] 3. Context modeling:
[0253] The generated summary reads: "The user picks up their child at 5:30 PM on Fridays and usually sets the air conditioner to 26°C."
[0254] LLM validation logic: "The time of returning home must be associated with the schedule and the current time (15:00);
[0255] 4. Environmental perception:
[0256] The indoor temperature was measured at 30℃.
[0257] The scenario is categorized as "daily commuting," and an efficient integration path is enabled.
[0258] 5. Intentional reasoning:
[0259] Output command: "17:20 Start the home air conditioner to 26℃".
[0260] Step S204: Control the target device according to the device control intent.
[0261] After controlling the target device according to the stated device control intent, user feedback (such as whether the intent was correctly executed) and environmental change data (such as user comfort after temperature adjustment) can be collected. The gating function, loss function, and reward mechanism are iteratively optimized to improve dynamic adaptability.
[0262] This invention achieves the following breakthroughs in the vehicle-home interconnection scenario through the deep integration of four core technologies: dynamic intent reasoning, multimodal fusion verification, privacy-graded encryption, and edge computing optimization:
[0263] Security: A dynamic encrypted communication system is built across the three terminals of CarCloudHome to effectively defend against attacks such as GPS spoofing and voiceprint forgery, reducing privacy and compliance risks.
[0264] User experience: Seamless collaborative control reduces user operation steps and improves satisfaction by combining personalized intent understanding (such as automatically associating with user habits).
[0265] In another embodiment of this application, a device control apparatus is also provided, such as... Figure 8 As shown, it includes:
[0266] The receiving module 11 is used to receive multimodal input information from the first linkage space and determine the target device in the second linkage space that the user intends to control based on the multimodal input information.
[0267] The acquisition module 12 is used to acquire the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category;
[0268] Verification module 13 is used to verify whether the user has control rights over the target device based on the multimodal input information and the device verification method;
[0269] Control module 14 is used to control the target device based on the multimodal input information if the user has control authority over the target device.
[0270] In another embodiment of this application, an electronic device is also provided, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0271] Memory, used to store computer programs;
[0272] The processor, when executing a program stored in memory, implements the device control method described in any of the foregoing method embodiments.
[0273] The electronic device provided in this embodiment of the invention has a processor that executes a program stored in a memory, determines a target device based on multimodal input information, and verifies the user's control permissions based on the target device's device privacy category and device verification method. When the user has control permissions over the target device, the processor controls the target device based on the multimodal input information and performs hierarchical verification of the user's control permissions based on the target device's device privacy category. This ensures that the user's control permissions match the device privacy category of the target device to be controlled, avoids leakage of privacy content of the target device that does not match the user's control permissions, and improves the security and privacy protection level of vehicle-home interconnection.
[0274] The communication bus 1140 mentioned in the above-mentioned electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus 1140 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0275] The communication interface 1120 is used for communication between the above-mentioned electronic device and other devices.
[0276] The memory 1130 may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0277] The processor 1110 mentioned above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0278] In another embodiment of this application, a computer-readable storage medium is also provided, on which a program for a device control method is stored, wherein when the program for the device control method is executed by a processor, it implements the steps of the device control method described in any of the foregoing method embodiments.
[0279] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0280] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A device control method characterized by, include: The system receives multimodal input information from a first linkage space. Based on this multimodal input information, it determines a target device for user intent control in a second linkage space. The target device is determined based on multimodal fusion information obtained by performing multimodal fusion processing on the multimodal input information. This multimodal fusion information is determined by evaluating the quality of each modal input information using an energy function, resulting in an energy uncertainty assessment result for each modality. The energy uncertainty assessment result is for each modal input information. Based on the modal input information, preset modal mean information, and a preset reconstruction error energy function, a first uncertainty assessment result is determined. Based on the feature representation of the modal input information, preset positive sample features, preset negative sample features, and a preset contrastive learning energy function, a second uncertainty assessment result is determined. Finally, the modal input information is input into a preset probability model energy function to determine a third uncertainty assessment result. Determined based on the first uncertainty assessment result, the second uncertainty assessment result, and the third uncertainty assessment result; Obtain the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category; Based on the multimodal input information and the device verification method, verify whether the user has control authority over the target device; If the user has control over the target device, the target device is controlled based on the multimodal input information.
2. The device control method according to claim 1, characterized by, Obtaining the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category includes: Obtain the core data type of the device associated with the target device; In the preset correspondence between device core data types and privacy categories, determine the device privacy category corresponding to the device core data type; In the preset correspondence between privacy categories and verification methods, determine the device verification method corresponding to the device privacy category.
3. The device control method according to claim 2, wherein In the preset correspondence between privacy categories and verification methods, the device verification method corresponding to the device privacy category is determined, including: If the device privacy category is high privacy category, the device verification method is a three-modal verification method; If the device privacy category is medium privacy category, the device verification method is bimodal verification method; If the device privacy category is low privacy category, the device verification method is single-modal verification method.
4. The device control method according to claim 1, wherein Verifying whether the user has control authority over the target device based on the multimodal input information and the device verification method includes: If the target device's device privacy category is high privacy, the user's face image, voice information, and location information are determined based on the multimodal input information. Trimodal verification is performed based on the face image, voice information, and location information. If the verification passes, it is determined that the user has control rights over the target device. If the target device's privacy category is medium privacy category, the user's voice information and location information are determined based on the multimodal input information. Bimodal verification is performed based on the voice information and location information. If the verification is successful, it is determined that the user has control rights over the target device. If the target device's privacy category is low privacy, the user's location information is determined based on the multimodal input information, and a single-modal verification is performed based on the location information. If the verification passes, it is determined that the user has control rights over the target device.
5. The equipment control method according to claim 1, characterized in that, Controlling the target device based on the multimodal input information includes: The multimodal input information is subjected to multimodal fusion processing to obtain multimodal fused information; Obtain a context summary of the multimodal input information; The device control intent is determined based on the multimodal fusion information and the context summary; Control the target device according to the stated device control intent.
6. The equipment control method according to claim 5, characterized in that, The multimodal input information is subjected to multimodal fusion processing to obtain multimodal fused information, including: The quality of each modal input information in the multimodal input information is evaluated using an energy function to obtain the energy uncertainty evaluation results for each mode; Based on the energy uncertainty assessment results of each mode, the multimodal input information is dynamically fused to obtain the multimodal fused information.
7. The equipment control method according to claim 6, characterized in that, Based on the energy uncertainty assessment results of each mode, dynamic multimodal fusion is performed on the multimodal input information to obtain the multimodal fused information, including: Obtain dynamic fusion weights for multiple modalities; If the energy uncertainty assessment result of any mode is lower than a preset threshold, the historical preference weight of that mode is obtained, the dynamic fusion weight of that mode is replaced by the historical preference weight, and the dynamic fusion weight of other modes is adjusted accordingly. The multimodal input information is fused based on the historical preference weights and the dynamic fusion weights of other modalities to obtain the multimodal fusion information.
8. The equipment control method according to claim 7, characterized in that, Obtain dynamic fusion weights for multiple modalities, including: Feature extraction is performed on the multimodal input information to obtain modal features of multiple modalities; Obtain the historical context of multiple modalities; Modal features and historical context of multiple modalities are input into a preset gating function to obtain modal-level weights of multiple modalities; Based on a preset resource-aware loss function, the modality-level weights of multiple modalities are dynamically adjusted to obtain dynamic fusion weights for multiple modalities.
9. The equipment control method according to claim 8, characterized in that, Modal features and historical context of multiple modalities are input into a preset gating function to obtain modal-level weights for multiple modalities, including: Modal features and historical context from multiple modalities are concatenated to obtain a joint vector; Obtain the degree of influence of the joint vector on the dynamic weights of each mode and the activation threshold of the gating function; The joint vector, the degree of influence, and the activation threshold are input into a gating network so that the gating network maps the joint vector into dynamic weights for each modality. The dynamic weights of each mode are normalized to obtain the modal-level weights of multiple modes.
10. The equipment control method according to claim 8, characterized in that, Based on a preset resource-aware loss function, the modality-level weights of multiple modalities are dynamically adjusted to obtain dynamic fusion weights for multiple modalities, including: The task loss value is determined based on the modal-level weights of multiple modalities; The instruction complexity is determined based on the multimodal input information; Obtain the complexity balance coefficient and resource complexity loss value corresponding to the instruction complexity; The resource-aware loss value is determined based on the task loss value, complexity balance coefficient, and resource complexity loss value. Based on the resource perception loss value, the modal-level weights of multiple modalities are dynamically adjusted to obtain the dynamic fusion weights of multiple modalities.
11. The equipment control method according to claim 6, characterized in that, Determining the device control intent based on the multimodal fusion information and the context summary includes: Obtain the context summary corresponding to the multimodal fusion information; The multimodal fusion information and the context summary are input into a preset intent prediction model, so that the intent prediction model outputs the device control intent corresponding to the multimodal fusion information and the context summary based on a multidimensional reward strategy and a dynamic KL divergence adjustment strategy.
12. A device control apparatus, characterized in that, include: A receiving module is configured to receive multimodal input information from a first linkage space, determine a target device for user intent control in a second linkage space based on the multimodal input information, wherein the target device is determined based on multimodal fusion information obtained by multimodal fusion processing of the multimodal input information, the multimodal fusion information being determined by evaluating the quality of each modal input information in the multimodal input information using an energy function, and obtaining an energy uncertainty assessment result for each modality, wherein the energy uncertainty assessment result is for the modal input information of each modality, and a first uncertainty assessment result is determined based on the modal input information, preset modal mean information, and preset reconstruction error energy function; a second uncertainty assessment result is determined based on the feature representation of the modal input information, preset positive sample features, preset negative sample features, and preset contrastive learning energy function; and a third uncertainty assessment result is determined by inputting the modal input information into a preset probability model energy function. Determined based on the first uncertainty assessment result, the second uncertainty assessment result, and the third uncertainty assessment result; The acquisition module is used to acquire the device privacy category corresponding to the target device and the device verification method corresponding to the device privacy category; The verification module is used to verify whether the user has control authority over the target device based on the multimodal input information and the device verification method. The control module is used to control the target device based on the multimodal input information if the user has control authority over the target device.
13. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the device control method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for a device control method, which, when executed by a processor, implements the steps of the device control method according to any one of claims 1-11.
Citation Information
Patent Citations
Authority control method and apparatus for household appliances, computer device and storage medium
CN110442033A
Equipment control method and device, electronic device and storage medium
CN116088331A
Multi-modal feature dynamic fusion method and system based on uncertainty estimation
CN120257217A
Multi-modal intention recognition method and system based on consistency and difference decoupling
CN120654179A