A multi-source data interaction control system and method for intelligent display
By constructing modal connectivity relationships and supplementing response models, the problem of ineffective user response to multimodal interactions on smart displays was solved, thereby improving interaction efficiency and user experience.
Patent Information
- Application Number
- CN202510615582.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Existing smart displays fail to respond effectively to user actions in multimodal interaction modes, resulting in decreased interaction efficiency and experience, especially when users are unfamiliar with the interaction methods, which can easily lead to erroneous operations.
By filtering and analyzing non-single-modal interactive events of smart displays, modal connectivity relationships and supplementary response models are constructed to generate prompt response content to improve the smoothness of interaction.
It enables accurate analysis and effective prompts for user operations under multimodal interaction, improving the interaction efficiency and user experience of smart displays.
Smart Images

Figure CN120544557B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interactive control technology, specifically a multi-source data interactive control system and method for smart displays. Background Technology
[0002] With the continuous advancement of technology, smart displays are no longer limited to displaying and playing images. They can also enable voice interaction or smart home interaction in certain scenarios, as well as multimodal interaction methods in complex scenarios, such as voice and touch interaction. However, the coexistence of multiple interaction methods can lead to issues with interaction coordination. For example, after a user operates using a voice command, the display may not respond effectively to the user's desired response, requiring manual adjustment of some parameters or further touch operations. When some users are not familiar with the smart interaction method and do not know how to use it, incorrect operations may occur, failing to effectively achieve the interaction purpose and affecting the efficiency and experience of the interaction. Summary of the Invention
[0003] The purpose of this invention is to provide a multi-source data interaction control system and method for smart displays, so as to solve the problems raised in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a multi-source data interaction control method for a smart display, the method comprising the following steps:
[0005] Step S100: Filter the smart display records of interactive operation events that are not single-modal. Interactive operation events refer to events in which the smart display is woken up by voice data to perform interactive operations. Interactive operation events record voice content, user operation content and smart display response content. Non-single-modal means that when the recorded smart display response content is the same, the modal interaction methods in each corresponding interactive operation event are different.
[0006] Step S200: Mark the interaction operation event when the number of recorded modal interaction methods is 1 for each type of response content as the first interaction event, extract the voice content in the first interaction event as the first voice content of the corresponding response content type; analyze and determine the target voice content based on the first voice content; and mark the interaction operation event corresponding to the target voice content as the target interaction event;
[0007] Step S300: Extract the remaining interactive events corresponding to the same type of response content, excluding the target interactive event, as reference interactive events; and construct the modal connection relationship of the reference interactive events;
[0008] Step S400: Based on the modal connectivity relationship, extract the voice content of the corresponding interactive event records and generate a supplementary response model corresponding to each modal connectivity relationship;
[0009] Step S500: Based on the supplementary response model, acquire the real-time captured voice content. If the smart display fails to trigger the response content within a preset time period, analyze the real-time voice content based on the supplementary response model and provide prompt response content.
[0010] Furthermore, determining the target speech content based on the analysis of the first speech content includes the following specific steps:
[0011] Record each first interactive event of the same type of response content and bind it with the corresponding first voice content; sort the first interactive events in ascending order according to the duration of the first voice content; extract the duration difference L1 between the first voice content recorded by the first and last first interactive events in the sequence;
[0012] If the duration difference L1 is less than or equal to the duration difference threshold L0, then extract the time length T between the initial reception time of the voice content and the display time of the response content for each first interactive event in the sequence, and calculate the duration dispersion Q of the first interactive event, Q={[∑(T-T0)} 2 ] / M} 1 / 2 Where M represents the number of first interactive events in the sequence; T0 represents the average recording time of the first interactive event; if the duration dispersion Q is less than or equal to the dispersion threshold Q0, all first interactive events in the output sequence are the target interactive events; if the duration dispersion Q is greater than the dispersion threshold Q0, the first interactive event corresponding to T less than T0 is selected as the target interactive event.
[0013] If the duration difference L1 is greater than the duration difference threshold L0, calculate the duration difference between each first interaction event and the first first interaction event in sequence, and select the first interaction event whose duration difference is less than the duration difference threshold as the target interaction event;
[0014] The voice content corresponding to the target interactive event is used as the target voice content.
[0015] Analyzing the target interaction events demonstrates that voice commands successfully wake up the smart display for interactive operations; it ensures that valid evaluation criteria are recorded in the successful wake-up interaction events.
[0016] Furthermore, constructing the modal connectivity relationships for the corresponding interactive events includes the following steps:
[0017] Extract the user operation content of the comparison interaction event. User operation content refers to control modalities other than user voice control modalities.
[0018] The monitoring interval is generated by taking the time when the response content is displayed in the comparison interaction event as the end time and the time when the user triggers the voice control modality as the start time. All control modalities within the monitoring interval are extracted, and the modal connection relationship of each comparison interaction event is generated by sequentially connecting the voice control modality as the initial modality of the modal connection relationship and the adjacent other control modalities in chronological order.
[0019] Furthermore, step S400 includes the following specific steps:
[0020] Step S410: Each type of response content contains at least one type of modal connection relationship. Extract the speech content in the corresponding interactive event of each modal connection relationship as the reference speech content. Calculate the similarity f1 between the reference speech content and the target speech content in the target interactive event corresponding to the same type of response content. Traverse and search all the reference interactive events of the same modal connection relationship record, calculate the similarity respectively, and generate the speech similarity set F of the corresponding modal connection relationship.
[0021] Step S420: Select the maximum value fmax and the minimum value fmin in the speech similarity set F, and construct a supplementary response model Y for the speech features corresponding to the modal connectivity relationship of each type of response content record; Y = (fmax - fmin) / fmax.
[0022] Analyzing supplementary response models helps reduce the inconvenience caused to users when voice control modal responses fail to be successful in a timely manner. By establishing supplementary response models, voice content analysis can be effectively achieved under the modal connection relationship of different response content, thereby enabling effective analysis of the further responses required by smart displays.
[0023] Furthermore, step S500 includes the following specific steps:
[0024] The real-time voice content is obtained and substituted into several supplementary response models stored in the smart display. The modal connectivity relationship corresponding to the supplementary response model is extracted from the real-time voice content. Satisfying the supplementary response model means that the similarity between the real-time voice content and the target voice content recorded by all models is calculated and replaced with the minimum value fmin in the supplementary response model to calculate the real-time output value. When the real-time output value is less than or equal to the output value of the corresponding supplementary response model, it is said to be satisfied.
[0025] When the real-time voice content satisfies a unique supplementary response model, the response content corresponding to the modal connection relationship of the supplementary response model record is extracted as the prompt response content.
[0026] When the supplementary response model satisfied by real-time voice content is not unique, the response content corresponding to the supplementary response model is extracted as the valid response content, along with the target interaction events in the historical records corresponding to the valid response content. The user screen after the voice content is transmitted is obtained for each type of valid response content and the corresponding target interaction event record. The user screen is obtained by the smart display based on the monitoring device under preset access permissions.
[0027] Extract the real-time user screen after the real-time voice content is emitted and before responding to other control modalities, and calculate the user behavior similarity Z with the target interaction events recorded by each effective response content. Based on the user behavior similarity and the real-time output value Y0 of the supplementary response model in which the corresponding effective response content is located, calculate the evaluation coefficient p of the modal connection relationship corresponding to each type of effective response content, p = k1*(1 / Z) + k2*Y0; where k1 and k2 represent the corresponding reference coefficients.
[0028] The effective response content corresponding to the minimum value of the evaluation coefficient calculated based on the modal connectivity relationship is selected as the prompt response content. The prompt response content refers to the content that the smart display analyzes and judges and prompts when it fails to effectively trigger a response based on the user's voice control modality, thereby improving the smoothness of smart device use caused by the addition of multimodal interaction methods.
[0029] A multi-source data interaction control system for a smart display, the system includes an interactive operation event filtering module, a first interactive event marking module, a target interactive event analysis module, a modal connection relationship construction module, a supplementary response model analysis module, and a prompt response content warning module;
[0030] The interactive operation event filtering module is used to filter interactive operation events recorded by the smart display that perform non-single-modal actions;
[0031] The first interactive event marking module is used to mark the interactive operation event when the number of modal interaction methods corresponding to each type of response content is 1 as the first interactive event;
[0032] The target interaction event analysis module is used to mark the interactive operation events corresponding to the target voice content as target interaction events;
[0033] The modal connection relationship building module is used to build modal connection relationships for interactive events;
[0034] The supplementary response model analysis module is used to generate supplementary response models corresponding to each modal connectivity relationship;
[0035] The prompt response content warning module is used to analyze real-time voice content based on the supplementary response model and provide prompt response content.
[0036] Furthermore, the target interaction event analysis module includes an interaction event sorting unit, a duration dispersion calculation unit, and a target interaction event determination unit;
[0037] The interactive event sorting unit is used to bind each first interactive event and its corresponding first voice content to the same type of response content; and to sort the first interactive events from smallest to largest according to the duration of the first voice content.
[0038] The duration dispersion calculation unit is used to calculate the duration dispersion of the first interactive event;
[0039] The target interaction event determination unit is used to determine the first interaction event that meets the requirements as the target interaction event.
[0040] Furthermore, the supplementary response model analysis module includes a speech similarity set generation unit and a supplementary response model establishment unit;
[0041] The speech similarity set generation unit is used to extract the speech content in the corresponding interactive event of each modal connection relationship as the reference speech content, calculate the similarity between the reference speech content and the target speech content in the target interactive event corresponding to the same type of response content, traverse and search all the reference interactive events of the same modal connection relationship record, calculate the similarity respectively, and generate the speech similarity set of the corresponding modal connection relationship.
[0042] The supplementary response model building unit is used to select the maximum and minimum values in the speech similarity set and construct a supplementary response model for the speech features corresponding to the modal connectivity relationship of each type of response content record.
[0043] Furthermore, the prompt response content warning module includes a real-time output value analysis unit, a behavior similarity analysis unit, an evaluation coefficient calculation unit, and a prompt response content output unit;
[0044] The real-time output value analysis unit is used to calculate the similarity between the real-time speech content and the target speech content recorded by all models, and replaces the minimum value in the supplementary response model to calculate the real-time output value.
[0045] The behavior similarity analysis unit is used to determine the relationship between the real-time output value and the output value of the supplementary response model;
[0046] The evaluation coefficient calculation unit is used to calculate the evaluation coefficient of the modal connectivity relationship corresponding to each type of valid response content;
[0047] The prompt response content output unit selects the valid response content corresponding to the minimum value of the evaluation coefficient calculated by modal connectivity relationship as the prompt response content.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] 1. This invention analyzes interactive operation events with irregular interaction modes based on historical records, and filters out voice content under effective interaction modes that can be used as evaluation criteria, thereby improving the accuracy of voice data acquisition and analysis during real-time interaction.
[0050] 2. Furthermore, it generates corresponding evaluation models for different interaction modalities, which can effectively classify and judge when users cannot touch the smart display in a timely manner under diverse interaction modalities. It also combines various sensor data of the smart display to further evaluate the interaction tendency for voice content in real time, and finally realizes effective reminders for untimely or unsuccessful touch response, thereby improving the smoothness of the smart display under modal interaction and enhancing the user interaction experience. Attached Figure Description
[0051] Figure 1 This is a schematic diagram of the structure of a multi-source data interaction control system for a smart display according to the present invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] Example: Figure 1 As shown, this invention provides a technical solution for a multi-source data interaction control system and method for smart displays. The multi-source data interaction control method for smart displays includes the following steps:
[0054] Step S100: Filter the smart display records of interactive operation events that are not single-modal. Interactive operation events refer to events in which the smart display is woken up by voice data to perform interactive operations. Interactive operation events record voice content, user operation content and smart display response content. Non-single-modal means that when the recorded smart display response content is the same, the modal interaction methods in each corresponding interactive operation event are different.
[0055] Step S200: Mark the interaction operation event when the number of recorded modal interaction methods is 1 for each type of response content as the first interaction event, extract the voice content in the first interaction event as the first voice content of the corresponding response content type; analyze and determine the target voice content based on the first voice content; and mark the interaction operation event corresponding to the target voice content as the target interaction event;
[0056] Step S300: Extract the remaining interactive events corresponding to the same type of response content, excluding the target interactive event, as reference interactive events; and construct the modal connection relationship of the reference interactive events;
[0057] Step S400: Based on the modal connectivity relationship, extract the voice content of the corresponding interactive event records and generate a supplementary response model corresponding to each modal connectivity relationship;
[0058] Step S500: Based on the supplementary response model, acquire the real-time captured voice content. If the smart display fails to trigger the response content within a preset time period, analyze the real-time voice content based on the supplementary response model and provide prompt response content.
[0059] Determining the target speech content based on the analysis of the first speech content includes the following specific steps:
[0060] Record each first interactive event of the same type of response content and bind it with the corresponding first voice content; sort the first interactive events in ascending order according to the duration of the first voice content; extract the duration difference L1 between the first voice content recorded by the first and last first interactive events in the sequence;
[0061] If the duration difference L1 is less than or equal to the duration difference threshold L0, then extract the time length T between the initial reception time of the voice content and the display time of the response content for each first interactive event in the sequence, and calculate the duration dispersion Q of the first interactive event, Q={[∑(T-T0)} 2 ] / M} 1 / 2 Where M represents the number of first interactive events in the sequence; T0 represents the average recording time of the first interactive event; if the duration dispersion Q is less than or equal to the dispersion threshold Q0, all first interactive events in the output sequence are the target interactive events; if the duration dispersion Q is greater than the dispersion threshold Q0, the first interactive event corresponding to T less than T0 is selected as the target interactive event.
[0062] If the duration difference L1 is greater than the duration difference threshold L0, calculate the duration difference between each first interaction event and the first first interaction event in sequence, and select the first interaction event whose duration difference is less than the duration difference threshold as the target interaction event;
[0063] The voice content corresponding to the target interactive event is used as the target voice content.
[0064] Analyzing the target interaction events demonstrates that voice commands successfully wake up the smart display for interactive operations; it ensures that valid evaluation criteria are recorded in the successful wake-up interaction events.
[0065] When the modal interaction mode is unique, it means that only the voice interaction mode is recorded. When there are multiple voice wake-ups and the modal interaction mode is unique, it means that the voice interaction mode has not effectively implemented the interaction operation, so this part of the interaction event should be removed.
[0066] Constructing modal connections between interactive events involves the following steps:
[0067] Extract the user operation content of the interaction event. User operation content refers to control modes other than user voice control mode, such as touch mode.
[0068] The monitoring interval is generated by taking the time when the response content is displayed in the comparison interaction event as the end time and the time when the user triggers the voice control modality as the start time. All control modalities within the monitoring interval are extracted, and the modal connection relationship of each comparison interaction event is generated by sequentially connecting the voice control modality as the initial modality of the modal connection relationship and the adjacent other control modalities in chronological order.
[0069] When the modal connection relationship in the interaction event only exists in the voice control modality, that is, when voice control is repeated multiple times, the modal connection relationship is an independent modal relationship.
[0070] Step S400 includes the following specific steps:
[0071] Step S410: Each type of response content contains at least one type of modal connection relationship. Extract the speech content in the corresponding interactive event of each modal connection relationship as the reference speech content. Calculate the similarity f1 between the reference speech content and the target speech content in the target interactive event corresponding to the same type of response content. Traverse and search all the reference interactive events of the same modal connection relationship record, calculate the similarity respectively, and generate the speech similarity set F of the corresponding modal connection relationship.
[0072] Step S420: Select the maximum value fmax and the minimum value fmin in the speech similarity set F, and construct a supplementary response model Y for the speech features corresponding to the modal connectivity relationship of each type of response content record; Y = (fmax - fmin) / fmax.
[0073] Analyzing supplementary response models helps reduce the inconvenience caused to users when voice control modal responses fail to be successful in a timely manner. By establishing supplementary response models, voice content analysis can be effectively achieved under the modal connection relationship of different response content, thereby enabling effective analysis of the further responses required by smart displays.
[0074] When the modal connectivity relationships of the same type of response content record include both independent modal relationships and modal connectivity relationships with multiple modal connections, the recorded speech content data may have differences; therefore, supplementary response models are established for different modal connectivity relationships.
[0075] Step S500 includes the following specific steps:
[0076] The real-time voice content is obtained and substituted into several supplementary response models stored in the smart display. The modal connectivity relationship corresponding to the supplementary response model is extracted from the real-time voice content. Satisfying the supplementary response model means that the similarity between the real-time voice content and the target voice content recorded by all models is calculated and replaced with the minimum value fmin in the supplementary response model to calculate the real-time output value. When the real-time output value is less than or equal to the output value of the corresponding supplementary response model, it is said to be satisfied.
[0077] When the real-time voice content satisfies a unique supplementary response model, the response content corresponding to the modal connection relationship of the supplementary response model record is extracted as the prompt response content.
[0078] When the supplementary response model satisfied by real-time voice content is not unique, the response content corresponding to the supplementary response model is extracted as the valid response content, along with the target interaction events in the historical records corresponding to the valid response content. The user screen after the voice content is transmitted is obtained for each type of valid response content and the corresponding target interaction event record. The user screen is obtained by the smart display based on the monitoring device under preset access permissions.
[0079] Extract the real-time user screen after the real-time voice content is emitted and before any other control modalities are responded to, and calculate the user behavior similarity Z with the target interaction events recorded by each effective response content. For example, in the AI fitness function of a smart display, the display turns on the camera to capture the user's body language for feature extraction, then performs feature quantification and standardization, and finally uses Euclidean distance to calculate similarity for data analysis. Based on the user behavior similarity and the real-time output value Y0 of the supplementary response model in which the corresponding effective response content is located, calculate the evaluation coefficient p of the modal connection relationship corresponding to each type of effective response content, p = k1*(1 / Z) + k2*Y0; where k1 and k2 represent the corresponding reference coefficients; set by the system.
[0080] The effective response content corresponding to the minimum value of the evaluation coefficient calculated based on the modal connectivity relationship is selected as the prompt response content. The prompt response content refers to the content that the smart display analyzes and judges and prompts when it fails to effectively trigger a response based on the user's voice control modality, thereby improving the smoothness of smart device use caused by the addition of multimodal interaction methods.
[0081] A multi-source data interaction control system for a smart display, the system includes an interactive operation event filtering module, a first interactive event marking module, a target interactive event analysis module, a modal connection relationship construction module, a supplementary response model analysis module, and a prompt response content warning module;
[0082] The interactive operation event filtering module is used to filter interactive operation events recorded by the smart display that perform non-single-modal actions;
[0083] The first interactive event marking module is used to mark the interactive operation event when the number of modal interaction methods corresponding to each type of response content is 1 as the first interactive event;
[0084] The target interaction event analysis module is used to mark the interactive operation events corresponding to the target voice content as target interaction events;
[0085] The modal connection relationship building module is used to build modal connection relationships for interactive events;
[0086] The supplementary response model analysis module is used to generate supplementary response models corresponding to each modal connectivity relationship;
[0087] The prompt response content warning module is used to analyze real-time voice content based on the supplementary response model and provide prompt response content.
[0088] The target interaction event analysis module includes an interaction event sorting unit, a duration dispersion calculation unit, and a target interaction event determination unit;
[0089] The interactive event sorting unit is used to bind each first interactive event and its corresponding first voice content to the same type of response content; and to sort the first interactive events from smallest to largest according to the duration of the first voice content.
[0090] The duration dispersion calculation unit is used to calculate the duration dispersion of the first interactive event;
[0091] The target interaction event determination unit is used to determine the first interaction event that meets the requirements as the target interaction event.
[0092] The supplementary response model analysis module includes a speech similarity set generation unit and a supplementary response model establishment unit;
[0093] The speech similarity set generation unit is used to extract the speech content in the corresponding interactive event of each modal connection relationship as the reference speech content, calculate the similarity between the reference speech content and the target speech content in the target interactive event corresponding to the same type of response content, traverse and search all the reference interactive events of the same modal connection relationship record, calculate the similarity respectively, and generate the speech similarity set of the corresponding modal connection relationship.
[0094] The supplementary response model building unit is used to select the maximum and minimum values in the speech similarity set and construct a supplementary response model for the speech features corresponding to the modal connectivity relationship of each type of response content record.
[0095] The alert response content warning module includes a real-time output value analysis unit, a behavior similarity analysis unit, an evaluation coefficient calculation unit, and an alert response content output unit;
[0096] The real-time output value analysis unit is used to calculate the similarity between the real-time speech content and the target speech content recorded by all models, and replaces the minimum value in the supplementary response model to calculate the real-time output value.
[0097] The behavior similarity analysis unit is used to determine the relationship between the real-time output value and the output value of the supplementary response model;
[0098] The evaluation coefficient calculation unit is used to calculate the evaluation coefficient of the modal connectivity relationship corresponding to each type of valid response content;
[0099] The prompt response content output unit selects the valid response content corresponding to the minimum value of the evaluation coefficient calculated by modal connectivity relationship as the prompt response content.
[0100] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A multi-source data interaction control method for a smart display, characterized in that: The method includes the following steps: Step S100: Filter the smart display records of interactive operation events that perform non-single modality. The interactive operation event refers to the event of waking up the smart display with voice data to perform interactive operation. The interactive operation event records the voice content, user operation content and smart display response content. The non-single modality means that when the recorded smart display response content is the same, the modal interaction mode in each corresponding interactive operation event is different. Step S200: Mark the interaction operation event when the number of recorded modal interaction methods is 1 for each type of response content as the first interaction event, extract the voice content in the first interaction event as the first voice content of the corresponding response content type; analyze and determine the target voice content based on the first voice content; and mark the interaction operation event corresponding to the target voice content as the target interaction event; Step S300: Extract the remaining interactive events corresponding to the same type of response content, excluding the target interactive event, as reference interactive events; and construct the modal connection relationship of the reference interactive events; Step S400: Based on the modal connectivity relationship, extract the voice content of the corresponding interactive event records and generate a supplementary response model corresponding to each modal connectivity relationship; Step S500: Based on the supplementary response model, acquire the real-time captured voice content. If the smart display fails to trigger the response content within a preset time period, analyze the real-time voice content based on the supplementary response model and provide prompt response content.
2. The multi-source data interaction control method for a smart display according to claim 1, characterized in that: The process of determining the target speech content based on the analysis of the first speech content includes the following specific steps: Record each first interactive event and bind it with the corresponding first voice content to the same type of response content; The first interactive events are sorted from smallest to largest according to the duration of the first voice content; Extract the duration difference L1 between the first and last first interactive events recorded in the sequence; If the duration difference L1 is less than or equal to the duration difference threshold L0, then extract the time length T between the initial reception time of the voice content and the display time of the response content for each first interactive event in the sequence, and calculate the duration dispersion Q of the first interactive event, Q={[∑(T-T0)} 2 ] / M} 1 / 2 ; Where M represents the number of first interaction events in the sequence; T0 represents the average recording time of the first interaction event; if the duration dispersion Q is less than or equal to the dispersion threshold Q0, all first interaction events in the output sequence are the target interaction events; If the duration dispersion Q is greater than the dispersion threshold Q0, the first interaction event corresponding to T being less than T0 is selected as the target interaction event; If the duration difference L1 is greater than the duration difference threshold L0, calculate the duration difference between each first interaction event and the first first interaction event in sequence, and select the first interaction event whose duration difference is less than the duration difference threshold as the target interaction event; The voice content corresponding to the target interactive event is used as the target voice content.
3. The multi-source data interaction control method for a smart display according to claim 1, characterized in that: The construction of the modal connection relationship of the comparative interaction events includes the following steps: Extract the user operation content of the comparison interaction event, where the user operation content refers to control modes other than the user's voice control mode; The monitoring interval is generated by taking the time when the response content is displayed in the comparison interaction event as the end time and the time when the user triggers the voice control modality as the start time. All control modalities within the monitoring interval are extracted, and the modal connection relationship of each comparison interaction event is generated by sequentially connecting the voice control modality as the initial modality of the modal connection relationship and the adjacent other control modalities in chronological order.
4. The multi-source data interaction control method for a smart display according to claim 2, characterized in that: Step S400 includes the following specific steps: Step S410: Each type of response content contains at least one type of modal connection relationship. Extract the speech content in the corresponding interactive event of each modal connection relationship as the reference speech content. Calculate the similarity f1 between the reference speech content and the target speech content in the target interactive event corresponding to the same type of response content. Traverse and search all the reference interactive events of the same modal connection relationship record, calculate the similarity respectively, and generate the speech similarity set F of the corresponding modal connection relationship. Step S420: Select the maximum value fmax and the minimum value fmin in the speech similarity set F, and construct a supplementary response model Y for the speech features corresponding to the modal connectivity relationship of each type of response content record; Y = (fmax - fmin) / fmax.
5. The multi-source data interaction control method for a smart display according to claim 4, characterized in that: Step S500 includes the following specific steps: The real-time voice content is obtained and substituted into several supplementary response models stored in the smart display. The modal connectivity relationship corresponding to the supplementary response model is extracted from the real-time voice content. The condition of satisfying the supplementary response model means that the real-time voice content is calculated to obtain the similarity with the target voice content recorded by all models, and the similarity is used to replace the minimum value fmin in the supplementary response model to calculate the real-time output value. When the real-time output value is less than or equal to the output value of the corresponding supplementary response model, it is said to be satisfied. When the real-time voice content satisfies a unique supplementary response model, the response content corresponding to the modal connection relationship of the supplementary response model record is extracted as the prompt response content. When the supplementary response model satisfied by the real-time voice content is not unique, the response content corresponding to the supplementary response model is extracted as the valid response content, along with the target interaction event in the historical record corresponding to the valid response content. The user screen after the voice content is sent is obtained for each type of valid response content and the corresponding target interaction event record. The user screen is obtained by the smart display based on the monitoring device under preset access permissions. Extract the real-time user screen after the real-time voice content is emitted and before responding to other control modalities, and calculate the user behavior similarity Z with the target interaction events recorded by each effective response content. Based on the user behavior similarity and the real-time output value Y0 of the supplementary response model in which the corresponding effective response content is located, calculate the evaluation coefficient p of the modal connection relationship corresponding to each type of effective response content, p = k1*(1 / Z) + k2*Y0; where k1 and k2 represent the corresponding reference coefficients. The effective response content corresponding to the minimum value of the evaluation coefficient calculated by modal connectivity relationship is selected as the prompt response content.
6. A multi-source data interaction control system for a smart display, as described in any one of claims 1-5, characterized in that: The system includes an interactive operation event filtering module, a first interactive event marking module, a target interactive event analysis module, a modal connection relationship construction module, a supplementary response model analysis module, and a prompt response content warning module; The interactive operation event filtering module is used to filter interactive operation events recorded by the smart display that perform non-single-modal actions. The first interactive event marking module is used to mark the interactive operation event when the number of recorded modal interaction methods is 1 for each type of response content as the first interactive event; The target interaction event analysis module is used to mark the interactive operation events corresponding to the target voice content as target interaction events; The modal connection relationship construction module is used to construct modal connection relationships for corresponding interactive events; The supplementary response model analysis module is used to generate a supplementary response model corresponding to each modal connection relationship; The prompt response content warning module is used to analyze real-time voice content based on the supplementary response model and provide prompt response content.
7. A multi-source data interaction control system for a smart display according to claim 6, characterized in that: The target interaction event analysis module includes an interaction event sorting unit, a duration dispersion calculation unit, and a target interaction event determination unit; The interactive event sorting unit is used to record each first interactive event of the same type of response content and bind it with the corresponding first voice content; The first interactive events are sorted from smallest to largest according to the duration of the first voice content; The duration dispersion calculation unit is used to calculate the duration dispersion of the first interactive event; The target interaction event determination unit is used to determine the first interaction event that meets the requirements as the target interaction event.
8. A multi-source data interaction control system for a smart display according to claim 6, characterized in that: The supplementary response model analysis module includes a speech similarity set generation unit and a supplementary response model establishment unit; The speech similarity set generation unit is used to extract the speech content in the corresponding interactive event of each modal connection relationship as the reference speech content, calculate the similarity between the reference speech content and the target speech content in the target interactive event corresponding to the same type of response content, traverse and search all the reference interactive events of the same modal connection relationship record, calculate the similarity respectively, and generate the speech similarity set of the corresponding modal connection relationship. The supplementary response model building unit is used to select the maximum and minimum values in the speech similarity set and construct a supplementary response model for the speech features corresponding to the modal connectivity relationship of each type of response content record.
9. A multi-source data interaction control system for a smart display according to claim 6, characterized in that: The prompt response content warning module includes a real-time output value analysis unit, a behavior similarity analysis unit, an evaluation coefficient calculation unit, and a prompt response content output unit; The real-time output value analysis unit is used to calculate the similarity between the real-time speech content and the target speech content recorded by all models, and replace the minimum value in the supplementary response model to calculate the real-time output value. The behavior similarity analysis unit is used to determine the relationship between the real-time output value and the output value of the supplementary response model; The evaluation coefficient calculation unit is used to calculate the evaluation coefficient of the modal connection relationship corresponding to each type of valid response content; The prompt response content output unit selects the valid response content corresponding to the minimum value of the evaluation coefficient calculated by modal connection relationship as the prompt response content.
Citation Information
Patent Citations
Semantic information description implementation method and device
CN116432657A
Context-aware control for smart devices
WO2019195799A1