Interaction control method and system based on scene space
Through the interactive control method based on scene space, user control intention information is obtained and target control schemes are generated, which solves the problem of rigid traditional interaction methods and improves interaction efficiency and user experience.
Patent Information
- Application Number
- CN202510132966.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional human-computer interaction method is stiff, and users need to confirm the interaction instructions multiple times, affecting the user experience and interaction efficiency.
The interactive control method based on scene space is adopted, by obtaining the user's control intention information, the interactive control instruction set is determined, and the model is determined using the pre-constructed control scheme, the target control scheme is generated, and the object to be controlled is then controlled.
It reduces the number of times the user performs interactive control operations, improves the user's interactive experience, and improves the interaction efficiency.
Smart Images

Figure CN120045068A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to an interactive control method and system based on scene space. Background Art
[0002] Traditional human-computer interaction methods are usually rather rigid, requiring users to confirm interaction commands multiple times throughout the entire interaction process, which affects the user's interaction experience and reduces interaction efficiency. Summary of the invention
[0003] In order to solve the above technical problems, the embodiments of the present application propose an interactive control method and system based on scene space, which can improve the user's interactive experience and improve the interaction efficiency.
[0004] In a first aspect, an embodiment of the present application provides an interactive control method based on a scene space, comprising:
[0005] Obtaining the user's control intention information for the object to be controlled;
[0006] Determining a corresponding interactive control instruction set based on the control intention information;
[0007] Based on the intention features of the control intention information, the control instruction set features of the interactive control instruction set, and the image features of the scene space image of the scene where the object to be controlled is located, a target control scheme is determined using a pre-built control scheme determination model, wherein the control scheme determination model is constructed by a target artificial intelligence model;
[0008] The object to be controlled is controlled according to the target control scheme.
[0009] Optionally, obtaining the user's control intention information for the object to be controlled includes:
[0010] Detecting the user's posture information through a sensor;
[0011] When it is determined that the posture information satisfies the control condition, the scene space image is acquired, the sight line information of the user is detected by a sensor, and sight line features of the sight line information and posture features of the posture information are extracted, wherein the scene space image includes the user and the object to be controlled;
[0012] In the scene space image, determining a first local area corresponding to the user and a second local area corresponding to the object to be controlled;
[0013] Extracting features from the scene space image based on the first local area to obtain a first local visual feature, and extracting features from the scene space image based on the second local area to obtain a second local visual feature;
[0014] The control intention information is generated using a first large language model based on the line of sight features, the posture features, the first local visual features, the second local visual features, and the image features of the scene space image.
[0015] Optionally, the generating the control intention information by using a first large language model based on the sight line feature, the posture feature, the first local visual feature, the second local visual feature and the image feature of the scene space image includes:
[0016] Converting the first local visual feature into a latent space to obtain a first latent feature, inputting the first latent feature into an encoder network for mapping to obtain a first position feature, and concatenating the first local visual feature and the first position feature to obtain a first region feature;
[0017] Converting the second local visual feature into a latent space to obtain a second latent feature, inputting the second latent feature into the encoder network for mapping to obtain a second position feature, and concatenating the second local visual feature and the second position feature to obtain a second region feature;
[0018] splicing the image feature, the first region feature and the second region feature to obtain a splicing feature;
[0019] The splicing feature, the sight feature and the posture feature are input into the first large language model to prompt the first large language model to generate the control intention information.
[0020] Optionally, the control scheme determines how the model is constructed, including:
[0021] Determine a plurality of large models sorted by target capability, wherein the target capability is used to indicate the accuracy of the control solution generated by the large model based on the intention feature, the control instruction set feature and the image feature, wherein the larger model is sorted higher, the lower the accuracy;
[0022] According to the order of the multiple large models, iteratively train the artificial intelligence model to be trained using the multiple large models;
[0023] The target artificial intelligence model is determined according to the artificial intelligence model after training, so as to construct the control scheme determination model using the target artificial intelligence model.
[0024] Optionally, the large model used in the i-th training round in the iterative training is the i-th large model among the multiple large models, i is a positive integer, and the iterative training of the artificial intelligence model to be trained using the multiple large models according to the order of the multiple large models includes:
[0025] For each training round in the iterative training, the large model of the current round is used to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model after training in the current round, wherein, if the current round is the first training round in the iterative training, then the artificial intelligence model that needs to be trained in the current round refers to the artificial intelligence model to be trained, otherwise it refers to the artificial intelligence model after training in the previous round.
[0026] Optionally, the using of the large model of the current round to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model trained in the current round includes:
[0027] Acquire a sample data set, wherein the sample data set includes a sample intention feature, a sample control instruction set feature, and a sample image feature;
[0028] Inputting the sample data set into the large model of the current round to obtain a first coding feature for characterizing a first control scheme;
[0029] Inputting the sample data set into the artificial intelligence model that needs to be trained in the current round to obtain a second coding feature for characterizing a second control scheme;
[0030] Based on the first encoding feature and the second encoding feature, the artificial intelligence model that needs to be trained in the current round is trained to obtain the artificial intelligence model after the current round of training.
[0031] Optionally, the training of the artificial intelligence model to be trained in the current round based on the first encoding feature and the second encoding feature to obtain the artificial intelligence model after the current round of training includes:
[0032] Constructing a loss function based on the first encoding feature and the second encoding feature;
[0033] The artificial intelligence model that needs to be trained in the current round is trained according to the loss function to obtain the artificial intelligence model after the current round of training.
[0034] Optionally, before determining the target control scheme based on the intention feature of the control intention information, the control instruction set feature of the interactive control instruction set, and the image feature of the scene space image of the scene where the object to be controlled is located, using a pre-built control scheme determination model, the method further includes:
[0035] Using the second language model, extracting features of the control intention information to obtain the intention features;
[0036] Using the third language model, extracting features of the interactive control instruction set to obtain features of the control instruction set;
[0037] The fourth language model is used to extract features of the scene space image to obtain the image features.
[0038] Optionally, after controlling the object to be controlled according to the target control scheme, the method further includes:
[0039] receiving evaluation information input by the user, wherein the evaluation information is used to represent the user's satisfaction with the target control solution;
[0040] In response to the evaluation information, the control scheme determination model is updated using the evaluation information.
[0041] In a second aspect, an embodiment of the present application provides an interactive control system based on a scene space, including:
[0042] An information acquisition module is used to obtain the user's control intention information for the object to be controlled;
[0043] An instruction set determination module, used to determine a corresponding interactive control instruction set based on the control intention information;
[0044] A control scheme generating module, for determining a target control scheme using a pre-built control scheme determining model based on the intention features of the control intention information, the control instruction set features of the interactive control instruction set, and the image features of the scene space image of the scene where the object to be controlled is located, wherein the control scheme determining model is constructed by a target artificial intelligence model;
[0045] A control module is used to control the object to be controlled according to the target control scheme.
[0046] In summary, the embodiments of the present application have at least the following beneficial effects:
[0047] According to the embodiment of the present application, the control intention information of the user for the object to be controlled is obtained; the corresponding interactive control instruction set is determined based on the control intention information; based on the intention characteristics of the control intention information, the control instruction set characteristics of the interactive control instruction set, and the image characteristics of the scene space image of the scene in which the object to be controlled is located, a target control scheme is determined using a pre-built control scheme determination model, wherein the control scheme determination model is constructed by a target artificial intelligence model; the object to be controlled is controlled according to the target control scheme, so that the control scheme can be automatically generated by using the artificial intelligence model in response to the control intention information, thereby reducing the number of times the user performs interactive control operations, improving the user's interactive experience, and improving the interaction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a flow chart of the interactive control method based on scene space provided in an embodiment of the present application;
[0049] Figure 2 It is a structural diagram of an interactive control system based on scene space provided in an embodiment of the present application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0051] In the description of the present application, the terms "first", "second", "third", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first", "second", "third", etc. may explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "multiple" is two or more. In the description of the present application, the term "including" and its variations are open inclusions, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "according to" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments".
[0052] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0053] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this application have the same meaning as those commonly understood by those skilled in the art. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood by specific circumstances.
[0054] First, see Figure 1 , shows a flow chart of an interactive control method based on scene space provided in an embodiment of the present application, the method includes steps S101-S104, which are as follows:
[0055] S101, obtaining the control intention information of the user for the object to be controlled.
[0056] In one example, the control intention information may be determined based on a text instruction, a voice instruction, and / or a gesture instruction input by a user.
[0057] S102: Determine a corresponding interactive control instruction set based on the control intention information.
[0058] In one example, step S102 may include: performing intent recognition on control intention information to parse user intention, such as using natural language processing technology or a machine learning model to analyze the control intention information input by the user to understand the real intention behind it, and mapping the parsed user intention to at least one corresponding control instruction to form an interactive control instruction set. For example, the instruction set can be pre-defined: according to the application field and system function, a complete set of interactive control instruction sets is pre-defined, so that the instructions in the interactive control instruction set can cover all possible operational requirements, and the identified user intention is mapped to at least one most suitable control instruction through a rule library or through a trained machine learning model.
[0059] In one example, the N interactive control instructions constituting the interactive control instruction set may be included in an interactive control instruction database. Specifically, the top N interactive control instructions with the highest similarity to the control intention information may be selected from the interactive control instruction database to form the interactive control instruction set.
[0060] S103, based on the intention characteristics of the control intention information, the control instruction set characteristics of the interactive control instruction set, and the image characteristics of the scene space image of the scene where the object to be controlled is located, a pre-built control scheme determination model is used to determine a target control scheme, wherein the control scheme determination model is constructed by a target artificial intelligence model.
[0061] In one example, the target artificial intelligence model / control scheme determination model can be a large language model deployed in the cloud with powerful multimodal data processing capabilities to facilitate the generation of accurate target control schemes.
[0062] It should be noted that artificial intelligence models are mathematical structures designed based on algorithms and statistical principles, aiming to simulate human cognitive functions such as learning, reasoning, and problem solving. These models process large amounts of data to identify patterns, extract features, and make predictions or decisions based on them. Depending on different task requirements and technical implementation methods, artificial intelligence models can be divided into many types, such as neural networks, support vector machines, decision trees, random forests, etc.
[0063] S104: Control the object to be controlled according to the target control scheme.
[0064] In one example, the object to be controlled may include a home appliance, which is deployed in a house, such as a residence, etc. For example, one or more home appliances are placed in a family residence or a public place.
[0065] In an optional implementation, the obtaining of the user's control intention information for the object to be controlled includes:
[0066] Detecting the user's posture information through a sensor;
[0067] When it is determined that the posture information satisfies the control condition, the scene space image is acquired, the sight line information of the user is detected by a sensor, and sight line features of the sight line information and posture features of the posture information are extracted, wherein the scene space image includes the user and the object to be controlled;
[0068] In the scene space image, determining a first local area corresponding to the user and a second local area corresponding to the object to be controlled;
[0069] Extracting features from the scene space image based on the first local area to obtain a first local visual feature, and extracting features from the scene space image based on the second local area to obtain a second local visual feature;
[0070] The control intention information is generated using a first large language model based on the line of sight features, the posture features, the first local visual features, the second local visual features, and the image features of the scene space image.
[0071] In one example, the sensor may include an image acquisition device such as a camera to detect the user's posture information and / or line of sight information.
[0072] In one example, the line of sight information may include: the distance between the user's eyes and the object to be controlled (for example, the distance may be the distance between the user's eyes and the center / bottom / top of the object to be controlled), the height difference between the user's eyes and the center of the object to be controlled, and / or the line of sight angle between the user's eyes and the object to be controlled.
[0073] In one example, the first large language model may adopt a multimodal large language model. The multimodal large language model has a better processing effect on multimodal input and can improve the accuracy of generating control intention information. The multimodal large language model can process input data corresponding to multiple modalities, for example, multiple modalities include text modality, visual modality, audio modality, and the combination of multiple modalities.
[0074] In one example, when it is detected that the gesture action indicated by the gesture information is a preset gesture action, it can be determined that the gesture information satisfies the control condition.
[0075] In one example, in addition to detecting the user's posture information through sensors to determine whether the control conditions are met, the user's voice commands can also be detected through sensors (such as microphones or other voice collection sensors). When the match between the detected voice commands and preset voice commands is higher than a preset threshold, it is determined that the control conditions are met.
[0076] In an optional implementation, the generating the control intention information using a first large language model based on the sight line feature, the posture feature, the first local visual feature, the second local visual feature and the image feature of the scene space image includes:
[0077] Converting the first local visual feature into a latent space to obtain a first latent feature, inputting the first latent feature into an encoder network for mapping to obtain a first position feature, and concatenating the first local visual feature and the first position feature to obtain a first region feature;
[0078] Converting the second local visual feature into a latent space to obtain a second latent feature, inputting the second latent feature into the encoder network for mapping to obtain a second position feature, and concatenating the second local visual feature and the second position feature to obtain a second region feature;
[0079] splicing the image feature, the first region feature and the second region feature to obtain a splicing feature;
[0080] The splicing feature, the sight feature and the posture feature are input into the first large language model to prompt the first large language model to generate the control intention information.
[0081] In this embodiment, local visual features can be used to characterize accurate and complete visual information of the corresponding local area, and the mapped position features can be used to characterize accurate and complete position information of the corresponding local area, so that the spliced regional features can accurately characterize the characteristic details of the corresponding local area, thereby making the final control intention information more accurate.
[0082] It should be noted that latent space is a concept used in machine learning and statistical modeling, especially in generative models (such as variational autoencoders, generative adversarial networks, etc.) and dimensionality reduction techniques. The latent space refers to an abstract space that is usually lower in dimension than the original data space. In this space, data points are represented as a set of latent variables, which can achieve compressed representation, that is, the latent space allows high-dimensional data to be mapped to a low-dimensional space, thereby achieving data compression while retaining important features. Exemplarily, the dimension indicated by the latent space can be one-dimensional.
[0083] In one example, the encoder network is a part of a neural network for mapping latent features to corresponding positional features, and the encoder network can achieve mapping by encoding.
[0084] In one example, the spliced feature, the sight feature, and the posture feature may be spliced into a prompt feature, so as to input the prompt feature into the first large language model.
[0085] In an optional implementation, the control scheme determines the construction mode of the model, including:
[0086] Determine a plurality of large models sorted by target capability, wherein the target capability is used to indicate the accuracy of the control solution generated by the large model based on the intention feature, the control instruction set feature and the image feature, wherein the larger model is sorted higher, the lower the accuracy;
[0087] According to the order of the multiple large models, iteratively train the artificial intelligence model to be trained using the multiple large models;
[0088] The target artificial intelligence model is determined according to the artificial intelligence model after training, so as to construct the control scheme determination model using the target artificial intelligence model.
[0089] In one example, an embodiment of the present application can be applied to a terminal. Generally, the computing power of a terminal is low, and it is not convenient to deploy a large model with high computing power requirements. Through this embodiment, a large model with higher performance can be used to iteratively train an artificial intelligence model with a smaller number of parameters, thereby constructing a control scheme determination model that is easy to deploy on the terminal.
[0090] In one example, the multiple large models may be obtained by fine-tuning and training different large language models using sample datasets respectively.
[0091] In an optional implementation, the large model used in the i-th training round in the iterative training is the i-th large model among the multiple large models, i is a positive integer, and the iterative training of the artificial intelligence model to be trained using the multiple large models according to the order of the multiple large models includes:
[0092] For each training round in the iterative training, the large model of the current round is used to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model after training in the current round, wherein, if the current round is the first training round in the iterative training, then the artificial intelligence model that needs to be trained in the current round refers to the artificial intelligence model to be trained, otherwise it refers to the artificial intelligence model after training in the previous round.
[0093] In one example, using the big model of the current round to train the artificial intelligence model that needs to be trained in the current round can include: inputting sample data into the big model of the current round and the artificial intelligence model that needs to be trained in the current round respectively, and then, based on the similarity between the output results of the big model of the current round and the artificial intelligence model that needs to be trained in the current round, training the artificial intelligence model that needs to be trained in the current round, wherein the similarity can be determined by calculating the feature similarity between the respective output results.
[0094] In an optional implementation, the method of using the large model of the current round to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model trained in the current round includes:
[0095] Acquire a sample data set, wherein the sample data set includes a sample intention feature, a sample control instruction set feature, and a sample image feature;
[0096] Inputting the sample data set into the large model of the current round to obtain a first coding feature for characterizing a first control scheme;
[0097] Inputting the sample data set into the artificial intelligence model that needs to be trained in the current round to obtain a second coding feature for characterizing a second control scheme;
[0098] Based on the first encoding feature and the second encoding feature, the artificial intelligence model that needs to be trained in the current round is trained to obtain the artificial intelligence model after the current round of training.
[0099] In one example, based on the first coding feature and the second coding feature, training the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model after training in the current round can include: calculating the similarity between the first coding feature and the second coding feature, and using the similarity to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model after training in the current round.
[0100] In an optional implementation, the training of the artificial intelligence model to be trained in the current round based on the first encoding feature and the second encoding feature to obtain the artificial intelligence model after the current round of training includes:
[0101] Constructing a loss function based on the first encoding feature and the second encoding feature;
[0102] The artificial intelligence model that needs to be trained in the current round is trained according to the loss function to obtain the artificial intelligence model after the current round of training.
[0103] In one example, the loss function can be used to constrain the similarity between the first coding feature and the second coding feature, and then, based on the sample data set and the loss function, the artificial intelligence model that needs to be trained in the current round is incrementally trained to update the model parameters of the artificial intelligence model that needs to be trained in the current round, thereby obtaining the artificial intelligence model after the current round of training, wherein the loss function can also be used to constrain: the difference between the understanding of the input features of the artificial intelligence model that undergoes incremental training and the understanding of the input features of the artificial intelligence model before incremental training. Therefore, this embodiment can improve training efficiency through lightweight incremental training, and while effectively improving model performance, it also takes into account the problem of training cost.
[0104] In an optional implementation, the intention feature based on the control intention information, the control instruction set feature of the interactive control instruction set, and the image feature of the scene space image of the scene where the object to be controlled is located, using a pre-built control scheme determination model, before determining the target control scheme, the method further includes:
[0105] Using the second language model, extracting features of the control intention information to obtain the intention features;
[0106] Using the third language model, extracting features of the interactive control instruction set to obtain features of the control instruction set;
[0107] The fourth language model is used to extract features of the scene space image to obtain the image features.
[0108] In an optional implementation, after controlling the object to be controlled according to the target control scheme, the method further includes:
[0109] receiving evaluation information input by the user, wherein the evaluation information is used to represent the user's satisfaction with the target control solution;
[0110] In response to the evaluation information, the control scheme determination model is updated using the evaluation information.
[0111] On the second aspect, accordingly, the embodiments of the present application also provide an interactive control system based on scene space, which can implement all processes of the interactive control method based on scene space provided in the above embodiments.
[0112] See also Figure 2 , shows a schematic diagram of the structure of an interactive control system based on scene space provided in an embodiment of the present application, the interactive control system based on scene space includes:
[0113] The information acquisition module 201 is used to acquire the control intention information of the user for the object to be controlled;
[0114] An instruction set determination module 202, configured to determine a corresponding interactive control instruction set based on the control intention information;
[0115] A control scheme generating module 203 is used to determine a target control scheme using a pre-built control scheme determination model based on the intention features of the control intention information, the control instruction set features of the interactive control instruction set, and the image features of the scene space image of the scene where the object to be controlled is located, wherein the control scheme determination model is constructed by a target artificial intelligence model;
[0116] The control module 204 is used to control the object to be controlled according to the target control scheme.
[0117] In an optional implementation, the obtaining of the user's control intention information for the object to be controlled includes:
[0118] Detecting the user's posture information through a sensor;
[0119] When it is determined that the posture information satisfies the control condition, the scene space image is acquired, the sight line information of the user is detected by a sensor, and sight line features of the sight line information and posture features of the posture information are extracted, wherein the scene space image includes the user and the object to be controlled;
[0120] In the scene space image, determining a first local area corresponding to the user and a second local area corresponding to the object to be controlled;
[0121] Extracting features from the scene space image based on the first local area to obtain a first local visual feature, and extracting features from the scene space image based on the second local area to obtain a second local visual feature;
[0122] The control intention information is generated using a first large language model based on the line of sight features, the posture features, the first local visual features, the second local visual features, and the image features of the scene space image.
[0123] In an optional implementation, the generating the control intention information using a first large language model based on the sight line feature, the posture feature, the first local visual feature, the second local visual feature and the image feature of the scene space image includes:
[0124] Converting the first local visual feature into a latent space to obtain a first latent feature, inputting the first latent feature into an encoder network for mapping to obtain a first position feature, and concatenating the first local visual feature and the first position feature to obtain a first region feature;
[0125] Converting the second local visual feature into a latent space to obtain a second latent feature, inputting the second latent feature into the encoder network for mapping to obtain a second position feature, and concatenating the second local visual feature and the second position feature to obtain a second region feature;
[0126] splicing the image feature, the first region feature and the second region feature to obtain a splicing feature;
[0127] The splicing feature, the sight feature and the posture feature are input into the first large language model to prompt the first large language model to generate the control intention information.
[0128] In an optional implementation, the control scheme determines the construction mode of the model, including:
[0129] Determine a plurality of large models sorted by target capability, wherein the target capability is used to indicate the accuracy of the control solution generated by the large model based on the intention feature, the control instruction set feature and the image feature, wherein the larger model is sorted higher, the lower the accuracy;
[0130] According to the order of the multiple large models, iteratively train the artificial intelligence model to be trained using the multiple large models;
[0131] The target artificial intelligence model is determined according to the artificial intelligence model after training, so as to construct the control scheme determination model using the target artificial intelligence model.
[0132] In an optional implementation, the large model used in the i-th training round in the iterative training is the i-th large model among the multiple large models, i is a positive integer, and the iterative training of the artificial intelligence model to be trained using the multiple large models according to the order of the multiple large models includes:
[0133] For each training round in the iterative training, the large model of the current round is used to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model after training in the current round, wherein, if the current round is the first training round in the iterative training, then the artificial intelligence model that needs to be trained in the current round refers to the artificial intelligence model to be trained, otherwise it refers to the artificial intelligence model after training in the previous round.
[0134] In an optional implementation, the method of using the large model of the current round to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model trained in the current round includes:
[0135] Acquire a sample data set, wherein the sample data set includes a sample intention feature, a sample control instruction set feature, and a sample image feature;
[0136] Inputting the sample data set into the large model of the current round to obtain a first coding feature for characterizing a first control scheme;
[0137] Inputting the sample data set into the artificial intelligence model that needs to be trained in the current round to obtain a second coding feature for characterizing a second control scheme;
[0138] Based on the first encoding feature and the second encoding feature, the artificial intelligence model that needs to be trained in the current round is trained to obtain the artificial intelligence model after the current round of training.
[0139] In an optional implementation, the training of the artificial intelligence model to be trained in the current round based on the first encoding feature and the second encoding feature to obtain the artificial intelligence model after the current round of training includes:
[0140] Constructing a loss function based on the first encoding feature and the second encoding feature;
[0141] The artificial intelligence model that needs to be trained in the current round is trained according to the loss function to obtain the artificial intelligence model after the current round of training.
[0142] In an optional embodiment, the system further includes a feature extraction module, wherein the feature extraction module is used to:
[0143] Before determining the target control scheme using a pre-built control scheme determination model based on the intention feature of the control intention information, the control instruction set feature of the interactive control instruction set, and the image feature of the scene space image of the scene where the object to be controlled is located:
[0144] Using the second language model, extracting features of the control intention information to obtain the intention features;
[0145] Using the third language model, extracting features of the interactive control instruction set to obtain features of the control instruction set;
[0146] The fourth language model is used to extract features of the scene space image to obtain the image features.
[0147] In an optional embodiment, the system further includes:
[0148] An evaluation information receiving module, used for receiving evaluation information input by the user after the object to be controlled is controlled according to the target control scheme, wherein the evaluation information is used to represent the user's satisfaction with the target control scheme;
[0149] A model updating module is used to respond to the evaluation information and update the control scheme determination model using the evaluation information.
[0150] In summary, the embodiments of the present application have at least the following beneficial effects:
[0151] According to the embodiment of the present application, the control intention information of the user for the object to be controlled is obtained; the corresponding interactive control instruction set is determined based on the control intention information; based on the intention characteristics of the control intention information, the control instruction set characteristics of the interactive control instruction set, and the image characteristics of the scene space image of the scene in which the object to be controlled is located, a target control scheme is determined using a pre-built control scheme determination model, wherein the control scheme determination model is constructed by a target artificial intelligence model; the object to be controlled is controlled according to the target control scheme, so that the control scheme can be automatically generated by using the artificial intelligence model in response to the control intention information, thereby reducing the number of times the user performs interactive control operations, improving the user's interactive experience, and improving the interaction efficiency.
[0152] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary hardware platform, and of course it can also be implemented entirely by hardware. Based on such an understanding, all or part of the contribution of the technical solution of the present application to the background technology can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM (Read-Only Memory) / RAM (Random Access Memory), a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application or some parts of the embodiments.
[0153] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.
Claims
1. An interactive control method based on scene space, characterized in that: include: Obtaining the user's control intention information for the object to be controlled; Determining a corresponding interactive control instruction set based on the control intention information; Based on the intention features of the control intention information, the control instruction set features of the interactive control instruction set, and the image features of the scene space image of the scene where the object to be controlled is located, a target control scheme is determined using a pre-built control scheme determination model, wherein the control scheme determination model is constructed by a target artificial intelligence model; The object to be controlled is controlled according to the target control scheme.
2. The method according to claim 1, characterized in that The obtaining of the user's control intention information for the object to be controlled includes: Detecting the user's posture information through a sensor; When it is determined that the posture information satisfies the control condition, the scene space image is acquired, the sight line information of the user is detected by a sensor, and sight line features of the sight line information and posture features of the posture information are extracted, wherein the scene space image includes the user and the object to be controlled; In the scene space image, determining a first local area corresponding to the user and a second local area corresponding to the object to be controlled; Extracting features from the scene space image based on the first local area to obtain a first local visual feature, and extracting features from the scene space image based on the second local area to obtain a second local visual feature; The control intention information is generated using a first large language model based on the line of sight features, the posture features, the first local visual features, the second local visual features, and the image features of the scene space image.
3. The method according to claim 2, characterized in that The generating the control intention information by using a first large language model based on the sight line feature, the posture feature, the first local visual feature, the second local visual feature and the image feature of the scene space image includes: Converting the first local visual feature into a latent space to obtain a first latent feature, inputting the first latent feature into an encoder network for mapping to obtain a first position feature, and concatenating the first local visual feature and the first position feature to obtain a first region feature; Converting the second local visual feature into a latent space to obtain a second latent feature, inputting the second latent feature into the encoder network for mapping to obtain a second position feature, and concatenating the second local visual feature and the second position feature to obtain a second region feature; splicing the image feature, the first region feature and the second region feature to obtain a splicing feature; The splicing feature, the sight feature and the posture feature are input into the first large language model to prompt the first large language model to generate the control intention information.
4. The method according to claim 1, characterized in that The control scheme determines how the model is constructed, including: Determine a plurality of large models sorted by target capability, wherein the target capability is used to indicate the accuracy of the control solution generated by the large model based on the intention feature, the control instruction set feature and the image feature, wherein the larger model is sorted higher, the lower the accuracy; According to the order of the multiple large models, iteratively train the artificial intelligence model to be trained using the multiple large models; The target artificial intelligence model is determined according to the artificial intelligence model after training, so as to construct the control scheme determination model using the target artificial intelligence model.
5. The method according to claim 4, characterized in that The large model used in the i-th training round in the iterative training is the i-th large model among the multiple large models, i is a positive integer, and the multiple large models are used to iteratively train the artificial intelligence model to be trained according to the order of the multiple large models, including: For each training round in the iterative training, the large model of the current round is used to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model after training in the current round, wherein, if the current round is the first training round in the iterative training, then the artificial intelligence model that needs to be trained in the current round refers to the artificial intelligence model to be trained, otherwise it refers to the artificial intelligence model after training in the previous round.
6. The method according to claim 5, characterized in that The method of using the current round of large models to train the artificial intelligence model that needs to be trained in the current round to obtain the artificial intelligence model trained in the current round includes: Acquire a sample data set, wherein the sample data set includes a sample intention feature, a sample control instruction set feature, and a sample image feature; Inputting the sample data set into the large model of the current round to obtain a first coding feature for characterizing a first control scheme; Inputting the sample data set into the artificial intelligence model that needs to be trained in the current round to obtain a second coding feature for characterizing a second control scheme; Based on the first encoding feature and the second encoding feature, the artificial intelligence model that needs to be trained in the current round is trained to obtain the artificial intelligence model after the current round of training.
7. The method according to claim 6, characterized in that The step of training the artificial intelligence model that needs to be trained in the current round based on the first encoding feature and the second encoding feature to obtain the artificial intelligence model after the current round of training includes: Constructing a loss function based on the first encoding feature and the second encoding feature; The artificial intelligence model that needs to be trained in the current round is trained according to the loss function to obtain the artificial intelligence model after the current round of training.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: based on the intention feature of the control intention information, the control instruction set feature of the interactive control instruction set, and the image feature of the scene space image of the scene where the object to be controlled is located, using a pre-built control scheme determination model, before determining the target control scheme. Using the second language model, extracting features of the control intention information to obtain the intention features; Using the third language model, extracting features of the interactive control instruction set to obtain features of the control instruction set; The fourth language model is used to extract features of the scene space image to obtain the image features.
9. The method according to any one of claims 1 to 7, characterized in that: After controlling the object to be controlled according to the target control scheme, the method further includes: receiving evaluation information input by the user, wherein the evaluation information is used to represent the user's satisfaction with the target control solution; In response to the evaluation information, the control scheme determination model is updated using the evaluation information.
10. An interactive control system based on scene space, characterized in that: include: An information acquisition module is used to obtain the user's control intention information for the object to be controlled; An instruction set determination module, used to determine a corresponding interactive control instruction set based on the control intention information; A control scheme generating module, for determining a target control scheme using a pre-built control scheme determining model based on the intention features of the control intention information, the control instruction set features of the interactive control instruction set, and the image features of the scene space image of the scene where the object to be controlled is located, wherein the control scheme determining model is constructed by a target artificial intelligence model; A control module is used to control the object to be controlled according to the target control scheme.