Intelligent screen interaction method and system based on large model
Patent Information
- Application Number
- CN202610812045.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-06
- Publication Date
- 2026-08-28
AI Technical Summary
[0003]然而,现有智能家居屏的技术方案仍存在诸多局限,未能真正贴合居家封闭空间的独有特性、家庭多成员共处的复杂场景,以及居家生活中隐性、碎片化的核心需求,导致其智能化体验流于表面
[0043] Breaking through the limitations of existing active command triggering, this system achieves accurate recognition of non-standard commands such as self-talk, muttering, and complaining in home scenarios through non-standard command identification and implicit demand analysis. This aligns with the natural and casual behavioral habits of home life, significantly improving the level of intelligent interaction. Through spatial cognition, it enables the automatic division of virtual ownership spaces in the home and spatial context analysis, solving the problem of semantic differences of the same command in different spaces and achieving spatial adaptive response of commands.
Smart Images

Figure CN122654538A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart home screen interaction technology, specifically a smart screen interaction method and system based on a large model. Background Technology
[0002] With the rapid development of the smart home industry, smart home screens, as the central hub of home interaction, are gradually penetrating into various scenarios of home life. Their core functions have expanded from the initial device control and information display to intelligence and scenario-based applications.
[0003] However, existing smart home screen technologies still have many limitations, failing to truly align with the unique characteristics of enclosed home spaces, the complex scenarios of multiple family members living together, and the implicit and fragmented core needs of home life. This results in a superficial smart experience. For example, the interaction logic of existing smart home screens is still limited to active command triggering, relying on users to issue standard and complete control commands. They can only passively respond to explicit commands and cannot achieve seamless environmental optimization, which contradicts the "natural and casual" behavioral habits of home life. At the same time, existing technologies focus more on the control and linkage of the devices themselves, lacking a deep understanding of the attributes of home spaces. They fail to consider the differences in spatial functions of different areas of the home (living room, bedroom, dining room, etc.) and the semantic differences of the same command in different spatial contexts. This leads to a lack of targeted command responses, prone to false responses and response deviations, and unable to achieve adaptive interaction in spatial context.
[0004] In order to solve the above problems, this invention provides a smart screen interaction method and system based on a large model. Summary of the Invention
[0005] To address the problems of the aforementioned solutions, this invention provides a smart screen interaction method and system based on a large model.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] A smart screen interaction system based on a large model includes a command corpus, a multimodal acquisition module, an interaction analysis module, and an interaction control module.
[0008] The instruction corpus is used to store instruction data of various non-standard instructions in user home scenarios, and to label the instruction corpus with demand tags.
[0009] Furthermore, based on the voice data of the user's family, we assess whether there are linguistic differences between the user's family and the command corpus in the command corpus;
[0010] No action is taken when there are no language differences in the assessment;
[0011] When assessing language differences, the instruction corpus stored in the instruction corpus is optimized and supplemented based on speech data.
[0012] Furthermore, the instruction corpus stored in the instruction corpus is optimized and supplemented based on the speech data, including:
[0013] Based on the reasons for language differences, obtain the user's family voice data and command corpus that do not have corresponding supplementary data in the command corpus;
[0014] The speech data is used to assess whether the instruction corpus needs to be supplemented.
[0015] When the evaluation requires corpus supplementation, corresponding supplementary corpus is generated based on the reasons for language differences and the instruction corpus. The supplementary corpus is then associated with the instruction corpus and stored in the instruction corpus.
[0016] When the evaluation does not require corpus supplementation, the optimization and supplementation of the instruction corpus will not be performed.
[0017] The multimodal acquisition module is used to acquire data and obtain multimodal data.
[0018] The interaction analysis module is used to analyze multimodal data based on the command corpus to obtain interaction commands, identify the spatial interaction position corresponding to the interaction command based on the multimodal data, determine the target control command based on the spatial interaction position and the interaction command, and send the target control command to the interaction control module.
[0019] Furthermore, the multimodal data is analyzed based on the instruction corpus, including:
[0020] Standard command recognition is performed on multimodal data to obtain standard command recognition results, which include standard command control and unrecognized standard commands.
[0021] When the standard indicator identification result is standard command control, the identified standard command will be marked as an interactive command;
[0022] When the standard instruction recognition result is that no standard instruction was recognized, non-standard instruction recognition is performed on the multimodal data based on the instruction corpus to obtain the non-standard instruction recognition result. The non-standard instruction recognition result includes non-standard instruction control and unrecognized non-standard instructions. When the non-standard instruction recognition result is non-standard instruction control, the recognized non-standard instruction is marked as an interactive instruction. When the non-standard instruction recognition result is that no non-standard instruction was recognized, no corresponding processing is performed.
[0023] Furthermore, based on the instruction corpus, non-standard instruction recognition is performed on multimodal data, including:
[0024] Multimodal data is matched based on the instruction corpus and supplementary corpus corresponding to various non-standard instructions in the instruction corpus, and the non-standard instruction recognition result is determined based on the matching result.
[0025] Furthermore, based on multimodal data, the spatial interaction location corresponding to the interaction command is identified, including:
[0026] The home space is divided into several spatial areas. The location of the person corresponding to the interactive command is identified. The corresponding spatial area is located according to the location. The spatial area is marked as the spatial interaction location. The interactive commands of non-standard commands are optimized and adjusted according to the spatial interaction location.
[0027] Furthermore, the family space is divided into several spatial areas. The criteria for dividing the spatial areas are based on the regional attributes of the areas, which include the living room, dining room, kitchen, balcony, corridor, and bedroom.
[0028] Furthermore, the family space is divided into several spatial areas, including:
[0029] Identify control optimization devices and obtain control reference data between various locations within the home space and various control optimization devices;
[0030] Based on control reference data, a merging evaluation is performed on adjacent locations to obtain the merging evaluation results between each adjacent location. The merging evaluation results include merging and not merging.
[0031] Merge adjacent locations that are deemed to be merged in the merge evaluation to obtain each merged region; perform a merge evaluation on each merged region, merge adjacent merged regions that are deemed to be merged in the merge evaluation, and mark non-adjacent merged regions that are deemed to be merged in the merge evaluation as similar regions; mark each merged region as a spatial region.
[0032] Furthermore, the target control commands are determined based on the spatial interaction location and interaction commands, including:
[0033] Preset the initial weight value for each spatial region, adjust the initial weight value of the corresponding spatial region according to the spatial interaction position, and mark the initial weight value of each spatial region after adjustment as the weight value;
[0034] Calculate the weight coefficient for each spatial region based on the weight values;
[0035] Acquire control reference data for each spatial region, and perform priority analysis on the control parameters of the corresponding control equipment based on the control reference data and weighting coefficients of each spatial region to determine the target control command.
[0036] The interactive control module is used for interactive control, receiving target control commands, and adjusting control according to the target control commands.
[0037] Furthermore, it also includes an environmental interaction module, which is used to interact with the environment, acquire environmental data in real time, including home decoration style, wall reflectivity data, indoor light source spectrum, and time period sunshine data; perform visual optimization analysis based on environmental data to obtain screen adjustment parameters, and adjust the screen according to the screen adjustment parameters.
[0038] A smart screen interaction method based on a large model, the method comprising:
[0039] Data collection was conducted to obtain multimodal data;
[0040] The multimodal data is analyzed based on a pre-defined command corpus to obtain interactive commands. The spatial interaction location corresponding to the interactive command is identified based on the multimodal data. The target control command is determined based on the spatial interaction location and the interactive command.
[0041] Control adjustments are made according to the target control instructions.
[0042] Compared with the prior art, the beneficial effects of the present invention are:
[0043] Breaking through the limitations of existing active command triggering, this system achieves accurate recognition of non-standard commands such as self-talk, muttering, and complaining in home scenarios through non-standard command identification and implicit demand analysis. This aligns with the natural and casual behavioral habits of home life, significantly improving the level of intelligent interaction. Through spatial cognition, it enables the automatic division of virtual ownership spaces in the home and spatial context analysis, solving the problem of semantic differences of the same command in different spaces and achieving spatial adaptive response of commands. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a block diagram illustrating the principle of the present invention. Detailed Implementation
[0046] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0047] like Figure 1 As shown, an intelligent screen interaction system based on a large model includes a command corpus, a multimodal acquisition module, an interaction analysis module, and an interaction control module.
[0048] The instruction corpus is used to store instruction data of various non-standard instructions in user home scenarios, and to label the instruction corpus with demand tags.
[0049] The demand tag is used to indicate the control demand corresponding to the instruction corpus. For example, if the instruction corpus is "It's so stuffy", the demand tag is "ventilation demand"; if the instruction corpus is "lower the temperature", the demand tag is "lower the air conditioner setting temperature".
[0050] In one embodiment, the instruction corpus stored in the instruction corpus is configured by the platform based on the various control instructions that may exist in a home scenario, as well as the corpus of various non-standard instructions that the control instructions have, to form various instruction corpora and corresponding requirement tags.
[0051] In one embodiment, the speech data from the user's home is used to assess whether there are language differences between the user's home speech data and the instruction corpus in the instruction corpus. This is because the instruction corpus in the instruction corpus is generally set according to common speech, such as Mandarin. However, users may have language differences due to dialects, slang, English, etc. Therefore, in order to improve the recognition accuracy of the instruction corpus, it is necessary to conduct a language difference assessment and make corresponding optimizations and supplements when language differences exist. The speech data can directly identify whether there are differences between the speech data and the common language corresponding to the instruction corpus in the instruction corpus.
[0052] No action is taken when there are no language differences in the assessment;
[0053] When assessing language differences, the instruction corpus stored in the instruction corpus is optimized and supplemented based on speech data.
[0054] In one embodiment, the instruction corpus stored in the instruction corpus is optimized and supplemented based on the speech data. Each instruction corpus is collected according to the speech data, and the supplementary corpus corresponding to the speech data is associated with the instruction corpus and stored in the instruction corpus.
[0055] In one embodiment, the command corpus stored in the command corpus is optimized and supplemented based on voice data. Considering that in practical applications, users' families may use a mix of dialects, English, and Mandarin, and English may only be used for communication and learning, and not for voice control, in this embodiment, the above embodiments can be used to filter out or not generate corresponding supplementary corpus. The process is as follows:
[0056] Based on the reasons for language differences, the voice data of the user's family and the command corpus without corresponding supplementary data are obtained. This refers to the command corpus that has not generated supplementary data due to differences in voice data such as dialects, slang, and foreign languages. For example, if the command corpus A has generated supplementary data due to dialect differences, but the current voice difference is due to foreign languages, and the command corpus A has no supplementary data corresponding to the foreign language, it is considered that there is no corresponding supplementary data.
[0057] The speech data is used to assess whether the instruction corpus needs to be supplemented.
[0058] When the evaluation requires corpus supplementation, corresponding supplementary corpus is generated based on the reasons for the language differences and the instruction corpus. For example, the instruction corpus is represented by the language corresponding to the language differences to form supplementary corpus. The supplementary corpus is then associated with the instruction corpus and stored in the instruction corpus.
[0059] When the evaluation does not require corpus supplementation, the optimization and supplementation of the instruction corpus will not be performed.
[0060] In one embodiment, the need for supplementation of the instruction corpus is assessed based on voice data. The assessment criterion is whether the user family will narrate the instruction corpus according to the voice corresponding to the reason for the language difference. If narration is possible, the assessment criterion is met, and supplementation of the corpus is deemed necessary. Based on this, the assessment can be conducted based on the user family's voice data in various ways. For example, the assessment can be based on whether the voice data indicates that the instruction corpus is narrated according to the reason for the language difference, i.e., the assessment is based on the actual occurrence. Alternatively, the assessment can be based on the voice data to predict whether the instruction corpus will be narrated, such as predicting whether the user family narrates other instruction corpuses in the language difference.
[0061] The multimodal acquisition module is used to collect data and obtain multimodal data, including home environment data, user behavior data, and interaction data.
[0062] For example, a far-field microphone array is used, which supports far-field sound pickup up to 5 meters away. It has noise reduction and echo cancellation functions and can capture user standard commands, non-standard commands (talking to oneself, muttering, complaining, ellipses, family slang, etc.) and ambient sounds.
[0063] It uses a low-power, low-light camera that does not collect facial images or store human image data. It only captures behavioral and environmental visual features such as the layout of the home space, the position of the human body, light intensity, and wall reflections.
[0064] It integrates temperature and humidity sensors, CO2 sensors, odor sensors, and noise sensors to collect real-time data on enclosed home environments, air quality, and acoustic environments.
[0065] It connects to smart home devices, collects the operating status, energy consumption data, and consumable balance data of each device, and enables real-time perception of device status.
[0066] The interaction analysis module is used to analyze multimodal data based on the command corpus to obtain interaction commands, identify the spatial interaction position corresponding to the interaction command based on the multimodal data, determine the target control command based on the spatial interaction position and the interaction command, and send the target control command to the interaction control module.
[0067] In one embodiment, analyzing multimodal data based on an instruction corpus includes:
[0068] Standard command recognition is performed on multimodal data to obtain standard command recognition results, which include standard command control and unrecognized standard commands.
[0069] When the standard indicator identification result is standard command control, the identified standard command will be marked as an interactive command;
[0070] When the standard instruction recognition result is that no standard instruction was recognized, non-standard instruction recognition is performed on the multimodal data based on the instruction corpus to obtain the non-standard instruction recognition result. The non-standard instruction recognition result includes non-standard instruction control and unrecognized non-standard instructions. When the non-standard instruction recognition result is non-standard instruction control, the recognized non-standard instruction is marked as an interactive instruction. When the non-standard instruction recognition result is that no non-standard instruction was recognized, no corresponding processing is performed.
[0071] In one embodiment, standard instruction recognition is performed on multimodal data, and the recognition is performed according to various standard instruction recognition methods currently in use, such as recognition based on a preset standard instruction library.
[0072] In one embodiment, non-standard instructions are identified from multimodal data based on an instruction corpus. The instruction corpus and supplementary corpus of various non-standard instructions stored in the instruction corpus are matched. When a match is successful, the corresponding non-standard instruction is determined, and the non-standard instruction identification result is obtained.
[0073] In one embodiment, non-standard instruction recognition can be performed on multimodal data based on an instruction corpus. Semantic association and contextual completion functions can also be performed through a preset large model to achieve the corresponding non-standard instruction recognition.
[0074] In one embodiment, identifying the spatial interaction location corresponding to the interaction command based on multimodal data includes:
[0075] The home space is divided into several spatial zones, and the interactive commands are optimized and adjusted based on each spatial zone. That is, the actual control needs are determined according to the differences of the interactive command in the spatial zone, mainly targeting non-standard interactive commands.
[0076] The spatial area where the person who issued the interaction command is located is determined and marked as the spatial interaction location.
[0077] It mainly depends on the location of the person who issued the interactive command.
[0078] Data on human dwelling locations, door and window locations, and usage times are used to automatically divide virtual ownership spaces such as living rooms, bedrooms, dining rooms, and entryways using spatial clustering algorithms. Each space is pre-labeled with functional attributes (e.g., living room: leisure, entertaining guests; bedroom: rest, sleeping; dining room: dining). This is combined with human dwelling time (≥5 minutes in a certain area is considered the current activity space) and usage time (e.g., 22:00-7:00, with bedrooms being high-frequency activity spaces) to optimize space division accuracy. After division, the space division results and functional attribute labels are synchronously transmitted to the lightweight large model module and non-standard instruction processing module on the client side, providing support for subsequent contextual analysis and requirement reasoning.
[0079] S22: Spatial context analysis, connecting to the lightweight large model module on the client side, calling the spatial semantic analysis model built into the lightweight large model on the client side to identify the semantic differences of the same command in different virtual ownership spaces; specifically, it is implemented by associating and matching the command text (the content after speech-to-text conversion) with the functional attribute tags of the virtual ownership space, and combining common sense of home scenarios (such as "turn on the lights" corresponding to the main light in the living room and the bedside lamp in the bedroom) to determine the spatial adaptation response logic of the command; for example, if the user says "turn it down a little", if the current virtual ownership space is the living room, it is resolved as "turn down the living room air conditioner temperature / light brightness"; if it is the bedroom, it is resolved as "turn down the bedroom air conditioner temperature / bedside lamp brightness", avoiding mis-response of commands.
[0080] In one embodiment, determining the target control command based on the spatial interaction location and interaction command includes:
[0081] Simulation analysis is performed on each spatial region to determine the target control commands for the corresponding control devices based on different interactive commands in different spatial regions.
[0082] Match the corresponding target control command based on the spatial interaction location and interaction command.
[0083] In one embodiment, the family space is divided into several spatial areas, which can be based on spatial attributes, such as dividing it into living room, dining room, kitchen, balcony, corridor, and various rooms to form different spatial areas.
[0084] In one embodiment, the home space is divided into several spatial zones. To accommodate the detailed control needs of certain devices, these zones are divided according to the nature of each device, forming personalized spatial zones for each device, including:
[0085] Identify various control devices requiring detailed control analysis, such as air conditioners, fresh air systems, and window control devices. These can be selected based on user needs or the platform can preset a device list for subsequent matching. Corresponding control devices are marked as control optimization devices. Obtain control reference data between various locations within the home space and various control optimization devices. This data represents the impact of corresponding control devices on different control conditions at a given location, and is marked as control reference data. This data can be obtained from historical data or simulated using simulation technology to obtain the impact data of control devices at corresponding locations under different control conditions (control parameters, control commands).
[0086] Based on control reference data, a merging evaluation is performed on adjacent locations to obtain the merging evaluation results between each adjacent location. The merging evaluation results include merging and not merging.
[0087] The adjacent positions that are deemed to be merged in the merge evaluation are merged to obtain each merged region. For example, if there are four positions A, B, C, and D, where AB, BC, and BD are adjacent in sequence, and the merge evaluation result between AB is "merge", then AB is merged into a merged region. Since the merge evaluation result between BC is "merge", the merged region of C and AB is merged to obtain a new merged region ABC. The merge evaluation result between BD is "no merge", so D cannot be merged into this merged region. This process is repeated to obtain each merged region until each merged region has no adjacent positions or the merge evaluation result of the merged region is "merge", meaning that adjacent merged regions are also analyzed and merged.
[0088] A merger assessment is performed on each merged area. The merged areas that are deemed to be merged are marked as being of the same type. Since they are not adjacent, they cannot be merged, but marking them as being of the same type makes it easier to use the same control method. Each merged area is marked as a spatial region.
[0089] In one embodiment, adjacent locations are evaluated for merging based on control reference data. If the control reference data are the same, the merging evaluation result is merging; otherwise, merging is not performed.
[0090] In one embodiment, adjacent positions are evaluated by merging based on control reference data. The similarity between the corresponding control reference data is calculated by considering the influence of different control commands and parameters. Weight coefficients for different control situations are preset and calculated using a cosine similarity algorithm. When the similarity is greater than the preset value, the merging evaluation result is merged; otherwise, no merging is performed.
[0091] In one embodiment, simulation analysis is performed on each spatial region, and the optimal control parameters for the corresponding interactive commands within the control region are determined based on the control reference data of each spatial region, thereby forming the target control command.
[0092] In one embodiment, the optimal control parameters are the optimal control parameters for that spatial region.
[0093] In one embodiment, the optimal control parameter is the comprehensive optimal control parameter after measuring each spatial region. However, the regional weight of the spatial region is adjusted, the initial weight value of each spatial region is preset, the weight value of the spatial region is adjusted according to the region where the interactive command is located, the weight coefficient of each spatial region is determined, and the comprehensive optimal control parameter is determined according to the influence data of various control parameters in each spatial region.
[0094] In one embodiment, determining the target control command based on the spatial interaction location and interaction command includes:
[0095] The system presets initial weight values for each spatial area, which can be set based on the importance of the spatial area's attributes and control functions. Alternatively, users can set these values directly, or they can be set according to the proportion of activity time for personnel at corresponding times. The system adjusts the initial weight values of corresponding spatial areas based on the location of spatial interactions. For example, it presets the weight value that should be increased if the interaction command of a corresponding control function is located in that spatial area, and then adjusts it accordingly later. Alternatively, the system can set a uniform weight value and mark the initial weight values of each spatial area after the initial weight values are adjusted as the weight values.
[0096] Calculate the weight coefficient of each spatial region based on the weight values, that is, the ratio of the weight value to the whole.
[0097] The control reference data for each spatial region is obtained. Based on the control reference data and weight coefficients of each spatial region, the control parameters of the corresponding control equipment are prioritized to determine the target control command. The score of the spatial region can be determined by comparing it with the best state of each spatial region and calculating the priority value in combination with the weight coefficient. Alternatively, priority analysis can be performed based on other priority algorithms.
[0098] The interactive control module is used for interactive control, receiving target control commands, and adjusting control according to the target control commands.
[0099] In one embodiment, the system further includes an environmental interaction module, which interacts with the environment and acquires environmental data in real time. The environmental data includes home decoration style (preset parameters for minimalist, Chinese, European, etc.), wall reflection data (captured by a visual acquisition unit and the reflection intensity is calculated), indoor light source spectrum (acquired by an environmental acquisition unit), and time-period sunshine data (combined with real-time time and spatial orientation, and preset sunshine intensity parameters). Based on the environmental data, visual optimization analysis is performed to obtain screen adjustment parameters such as color shift, reflection suppression, blue light band segmented filtering, and screen contrast. The screen is then adjusted according to these parameters.
[0100] In one embodiment, visual optimization analysis based on environmental data can be performed using visual optimization algorithms, such as color shift correction algorithms. These algorithms employ scene-based dynamic color gamut calibration algorithms, which differ from general color shift correction algorithms. They can combine home decoration style parameters and indoor light source spectral data to dynamically calibrate the RGB three-color gain of the screen and solve color shift problems under different home light sources.
[0101] Reflection suppression algorithm: An adaptive ambient reflection suppression algorithm is adopted. By collecting data on the intensity of wall reflection, the brightness distribution of screen pixels is adjusted to reduce the pixel brightness in reflective areas and improve the contrast in non-reflective areas, thereby reducing the interference of home wall reflections on screen vision.
[0102] Blue light band segmented filtering algorithm + dynamic contrast adjustment algorithm: The two work together. The blue light filtering adopts a band-level adaptive algorithm to filter high-energy blue light (380-450nm) in segments according to the light intensity, avoiding image distortion caused by single filtering. The contrast adjustment adopts a scene-adaptive contrast algorithm, which dynamically adjusts the difference between brightness and darkness of the image based on the time of day, sunlight and light intensity, taking into account both eye protection and visual clarity.
[0103] For example, when the light intensity is ≥500 lux, the screen brightness is adjusted to 60%-80% and the color temperature is adjusted to 5000K-6500K to enhance reflection suppression; when the light intensity is <500 lux, the screen brightness is adjusted to 30%-50% and the color temperature is adjusted to 3000K-4500K to reduce blue light output; at the same time, according to the home scenario (dining, putting to sleep, entertaining guests, etc.), screen pop-ups and dynamic flashing images are restricted to achieve contextualized silent rendering of screen content (such as in the putting to sleep scenario, the screen automatically switches to dark screen mode, only displaying core information such as time, temperature and humidity).
[0104] In one embodiment, due to the controllability of smart homes, adjustments can be made by linking curtain control devices, lighting devices, etc.
[0105] A smart screen interaction method based on a large model, the method comprising:
[0106] Data collection was conducted to obtain multimodal data;
[0107] The multimodal data is analyzed based on a pre-defined command corpus to obtain interactive commands. The spatial interaction location corresponding to the interactive command is identified based on the multimodal data. The target control command is determined based on the spatial interaction location and the interactive command.
[0108] Control adjustments are made according to the target control instructions.
[0109] The above formulas are all numerical calculations after removing dimensions. The formulas are obtained by software simulation based on a large amount of data and are closest to the real situation. The preset parameters and preset thresholds in the formulas are set by those skilled in the art according to the actual situation or obtained by simulation based on a large amount of data.
[0110] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A smart screen interaction system based on a large model, characterized in that, It includes a command corpus, a multimodal acquisition module, an interactive analysis module, and an interactive control module; The instruction corpus is used to store instruction data of various non-standard instructions in user home scenarios, and to label the instruction corpus with demand tags; The multimodal acquisition module is used to acquire data and obtain multimodal data; The interaction analysis module is used to analyze multimodal data based on the command corpus to obtain interaction commands, identify the spatial interaction position corresponding to the interaction command based on the multimodal data, determine the target control command based on the spatial interaction position and the interaction command, and send the target control command to the interaction control module. The interactive control module is used for interactive control, receiving target control commands, and adjusting control according to the target control commands.
2. The intelligent screen interaction system based on a large model according to claim 1, characterized in that, Based on the user's home voice data, assess whether there are language differences between the user's home and the command corpus in the command corpus; No action is taken when there are no language differences in the assessment; When assessing language differences, the instruction corpus stored in the instruction corpus is optimized and supplemented based on speech data.
3. The intelligent screen interaction system based on a large model according to claim 2, characterized in that, The command corpus stored in the command corpus is optimized and supplemented based on speech data, including: Based on the reasons for language differences, obtain the user's family voice data and command corpus that do not have corresponding supplementary data in the command corpus; The speech data is used to assess whether the instruction corpus needs to be supplemented. When the evaluation requires corpus supplementation, corresponding supplementary corpus is generated based on the reasons for language differences and the instruction corpus. The supplementary corpus is then associated with the instruction corpus and stored in the instruction corpus. When the evaluation does not require corpus supplementation, the optimization and supplementation of the instruction corpus will not be performed.
4. The intelligent screen interaction system based on a large model according to claim 1, characterized in that, Analysis of multimodal data based on the instruction corpus, including: Standard command recognition is performed on multimodal data to obtain standard command recognition results, which include standard command control and unrecognized standard commands. When the standard indicator identification result is standard command control, the identified standard command will be marked as an interactive command; When the standard instruction recognition result is that no standard instruction was recognized, non-standard instruction recognition is performed on the multimodal data based on the instruction corpus to obtain the non-standard instruction recognition result. The non-standard instruction recognition result includes non-standard instruction control and unrecognized non-standard instructions. When the non-standard instruction recognition result is non-standard instruction control, the recognized non-standard instruction is marked as an interactive instruction. When the non-standard instruction recognition result is that no non-standard instruction was recognized, no corresponding processing is performed.
5. The intelligent screen interaction system based on a large model according to claim 4, characterized in that, Non-standard command recognition based on command corpus of multimodal data, including: Multimodal data is matched based on the instruction corpus and supplementary corpus corresponding to various non-standard instructions in the instruction corpus, and the non-standard instruction recognition result is determined based on the matching result.
6. The intelligent screen interaction system based on a large model according to claim 5, characterized in that, Based on multimodal data, the spatial interaction location corresponding to the interaction command is identified, including: The home space is divided into several spatial areas. The location of the person corresponding to the interactive command is identified. The corresponding spatial area is located according to the location. The spatial area is marked as the spatial interaction location. The interactive commands of non-standard commands are optimized and adjusted according to the spatial interaction location.
7. The intelligent screen interaction system based on a large model according to claim 6, characterized in that, Divide the home space into several spatial areas, including: Identify control optimization devices and obtain control reference data between various locations within the home space and various control optimization devices; Based on control reference data, a merging evaluation is performed on adjacent locations to obtain the merging evaluation results between each adjacent location. The merging evaluation results include merging and not merging. Merge adjacent locations that are deemed to be merged in the merge evaluation to obtain each merged region; perform a merge evaluation on each merged region, merge adjacent merged regions that are deemed to be merged in the merge evaluation, and mark non-adjacent merged regions that are deemed to be merged in the merge evaluation as similar regions; mark each merged region as a spatial region.
8. The intelligent screen interaction system based on a large model according to claim 1, characterized in that, Determine the target control commands based on the spatial interaction location and interaction commands, including: Preset the initial weight value for each spatial region, adjust the initial weight value of the corresponding spatial region according to the spatial interaction position, and mark the initial weight value of each spatial region after adjustment as the weight value; Calculate the weight coefficient for each spatial region based on the weight values; Acquire control reference data for each spatial region, and perform priority analysis on the control parameters of the corresponding control equipment based on the control reference data and weighting coefficients of each spatial region to determine the target control command.
9. The intelligent screen interaction system based on a large model according to claim 1, characterized in that, It also includes an environmental interaction module, which is used to interact with the environment and acquire environmental data in real time. The environmental data includes home decoration style, wall reflectivity data, indoor light source spectrum, and time period sunshine data. Based on the environmental data, visual optimization analysis is performed to obtain screen adjustment parameters, and screen adjustments are made according to the screen adjustment parameters.
10. A smart screen interaction method based on a large model, characterized in that, Applied to a large-model-based intelligent screen interaction system as described in any one of claims 1 to 9, the method includes: Data collection was conducted to obtain multimodal data; The multimodal data is analyzed based on a pre-defined command corpus to obtain interactive commands. The spatial interaction location corresponding to the interactive command is identified based on the multimodal data. The target control command is determined based on the spatial interaction location and the interactive command. Control adjustments are made according to the target control instructions.