Multi-source data interaction control system and method for intelligent display
By analyzing the multimodal interaction operation events of the smart display, building a model connection relationship and a supplementary response model, the problem of smart display untimely response under multimodal interaction is solved, and a more efficient interactive experience is achieved.
Patent Information
- Application Number
- CN202510615582.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Under multimodal interaction mode, smart displays fail to respond effectively to user needs, resulting in a decrease in interaction efficiency and experience, especially when users are not familiar with the interaction mode, it is easy to cause wrong operations.
By filtering and analyzing non-single modal interaction events for smart displays, building a modal connection relationship and supplementary response model, generating prompt response content to improve interaction fluency.
It improves the interaction accuracy and user experience under multimodal interaction, ensures that users can be effectively prompted to operate when voice control does not respond in time, and improves the interaction fluency of smart displays.
Smart Images

Figure CN120544557A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interactive control technology, and in particular to a multi-source data interactive control system and method for an intelligent display. Background Art
[0002] With the continuous advancement of technology, the application of smart displays is not only for the display and playback of images, but also for voice interaction or smart home interaction in certain scenarios; as well as multimodal interaction methods in complex scenarios, such as voice and touch interaction. However, the coexistence of multiple interaction methods will bring problems of interaction coordination. For example, after the user uses voice commands, the display fails to effectively respond to the content required by the user, and may need to manually adjust some parameters or perform further touch operations. When some users do not have a deep understanding of the smart interaction method and do not know how to use it, incorrect operations will occur, and the purpose of interaction cannot be effectively achieved, affecting the efficiency and experience of the interaction. Summary of the Invention
[0003] The object of the present invention is to provide a multi-source data interactive control system and method for an intelligent display to solve the problems raised in the prior art.
[0004] To achieve the above object, the present invention provides the following technical solution: a multi-source data interactive control method for an intelligent display, the method comprising the following steps:
[0005] Step S100: Filtering the smart display's records of interactive operation events that are not single-modal. An interactive operation event refers to an event that uses voice data to wake up the smart display for interactive operation. The interactive operation event records the voice content, user operation content, and smart display response content. Non-single-modal refers to recording the same smart display response content but with different modal interaction modes in each corresponding interactive operation event.
[0006] Step S200: Marking the interactive operation event corresponding to each type of response content when the number of recorded modal interaction modes is 1 as a first interactive event, extracting the voice content in the first interactive event as the first voice content corresponding to the response content type; determining the target voice content based on the first voice content analysis; and marking the interactive operation event corresponding to the target voice content as the target interactive event;
[0007] Step S300: extracting the remaining interactive operation events corresponding to the same type of response content except the target interactive event as reference interactive events; and constructing a modal connection relationship of the reference interactive events;
[0008] Step S400: extracting the speech content recorded for the interaction event based on the modal connection relationship, and generating a supplementary response model corresponding to each modal connection relationship;
[0009] Step S500: Based on the supplementary response model, the real-time captured voice content is obtained. If the smart display fails to successfully trigger the response content within the preset time period, the real-time voice content is analyzed based on the supplementary response model and a prompt response content is made.
[0010] Furthermore, determining the target speech content based on the first speech content analysis includes the following specific steps:
[0011] Bind each first interaction event recorded with the same type of response content to the corresponding first voice content; sort the first interaction events from small to large according to the duration of the first voice content; extract the duration difference L1 of the first voice content recorded between the first first interaction event and the last first interaction event in the sequence;
[0012] If the duration difference L1 is less than or equal to the duration difference threshold L0, then extract the time length T between the time when the smart display initially receives the voice content and the time when the response content is displayed for each first interaction event record in the sequence, and calculate the duration dispersion Q of the first interaction event, Q = {[∑(T-T0) 2 ] / M} 1 / 2 Where M represents the number of first interaction events in the sequence; T0 represents the average duration of the first interaction event records; if the duration dispersion Q is less than or equal to the dispersion threshold Q0, all first interaction events in the sequence are output as target interaction events; if the duration dispersion Q is greater than the dispersion threshold Q0, the first interaction event corresponding to T less than T0 is selected as the target interaction event;
[0013] If the duration difference L1 is greater than the duration difference threshold L0, the duration difference between each first interaction event and the first first interaction event is calculated in sequence, and the first interaction event whose duration difference is less than the duration difference threshold is selected as the target interaction event;
[0014] The voice content corresponding to the target interaction event is used as the target voice content.
[0015] Analyzing the target interaction event indicates that the voice has successfully awakened the smart display for interactive operation; ensuring that effective evaluation criteria are recorded in the operation event of successful awakening interaction.
[0016] Furthermore, constructing the modal connection relationship of the control interaction event includes the following steps:
[0017] Extracting user operation content of the control interaction event, where the user operation content refers to other control modes excluding the user voice control mode;
[0018] A monitoring interval is generated with the moment when the response content in the control interaction event is displayed as the end moment and the moment when the user triggers the voice control mode as the start moment. All control modes within the monitoring interval are extracted, and the modal connection relationship of each control interaction event is generated by connecting the voice control mode as the initial mode of the modal connection relationship and the other adjacent control modes in chronological order.
[0019] Furthermore, step S400 includes the following specific steps:
[0020] Step S410: Each type of response content contains at least one type of modal connection relationship. The speech content in the comparison interaction event corresponding to each modal connection relationship is extracted as the comparison speech content. The similarity f1 between the comparison speech content and the target speech content in the target interaction event corresponding to the same type of response content is calculated. All comparison interaction events recorded with the same modal connection relationship are traversed and the similarity is calculated for each event, generating a speech similarity set F corresponding to the modal connection relationship.
[0021] Step S420: Select the maximum value fmax and the minimum value fmin in the speech similarity set F, and construct a supplementary response model Y corresponding to the speech features of each type of response content recording modal connection relationship; Y = (fmax-fmin) / fmax.
[0022] The purpose of analyzing the supplementary response model is to reduce the inconvenience caused to users when the voice control modal response cannot be successful in a timely manner. By establishing a supplementary response model, the voice content analysis under the modal connection relationship corresponding to different response contents can be effectively realized, thereby realizing effective analysis of the further response required by the smart display.
[0023] Furthermore, step S500 includes the following specific steps:
[0024] Obtaining real-time speech content and substituting it into several supplementary response models stored in the intelligent display, extracting the real-time speech content that satisfies the modal connection relationship corresponding to the supplementary response model. Satisfying the supplementary response model means calculating the similarity between the real-time speech content and the target speech content recorded by all models, and replacing the minimum value fmin in the supplementary response model to obtain the real-time output value. When the real-time output value is less than or equal to the output value of the corresponding supplementary response model, it is considered satisfied.
[0025] When the real-time speech content satisfies a unique supplementary response model, extracting the response content corresponding to the modal connection relationship recorded by the supplementary response model as the prompt response content;
[0026] When the supplementary response model satisfied by the real-time voice content is not unique, the response content corresponding to the supplementary response model is extracted as the effective response content, and the target interaction event corresponding to the effective response content in the historical record is obtained. The user screen after the voice content of each type of effective response content corresponding to the target interaction event record is obtained. The user screen is obtained by the smart display based on the preset access permission of the monitoring device.
[0027] Extract the real-time user screen after the real-time voice content is sent and before responding to other control modes, and calculate the user behavior similarity Z with the target interaction event recorded by each valid response content. Based on the user behavior similarity and the real-time output value Y0 of the supplementary response model in which the corresponding valid response content is located, calculate the evaluation coefficient p of the modal connection relationship corresponding to each type of valid response content, p = k1*(1 / Z)+k2*Y0; where k1 and k2 represent corresponding reference coefficients;
[0028] The effective response content corresponding to the minimum value of the modal connection relationship evaluation coefficient is selected as the prompt response content. The prompt response content refers to the prompt content that the smart display analyzes and judges when the user fails to effectively trigger a response in the voice control mode, thereby improving the user's experience of the smart device due to the increase in multimodal interaction.
[0029] A multi-source data interaction control system for an intelligent display, the system comprising an interactive operation event screening module, a first interactive event marking module, a target interactive event analysis module, a modal connection relationship construction module, a supplementary response model analysis module, and a prompt response content warning module;
[0030] The interactive operation event screening module is used to screen interactive operation events recorded by the smart display that are not executed in a single mode;
[0031] The first interaction event marking module is used to mark the interaction operation event when the number of recorded modal interaction modes corresponding to each type of response content is 1 as the first interaction event;
[0032] The target interaction event analysis module is used to mark the interaction operation event corresponding to the target voice content as the target interaction event;
[0033] The modal connection relationship construction module is used to construct the modal connection relationship of the control interaction event;
[0034] The supplementary response model analysis module is used to generate a supplementary response model corresponding to each modal connection relationship;
[0035] The prompt response content warning module is used to analyze the real-time voice content based on the supplementary response model and make prompt response content.
[0036] Furthermore, the target interaction event analysis module includes an interaction event sorting unit, a duration dispersion calculation unit, and a target interaction event determination unit;
[0037] The interaction event sorting unit is used to bind each first interaction event recorded with the same type of response content to the corresponding first voice content; and sort the first interaction events from small to large according to the time length of the first voice content;
[0038] The duration dispersion calculation unit is used to calculate the duration dispersion of the first interaction event;
[0039] The target interaction event determining unit is configured to determine a first interaction event that meets the requirements as a target interaction event.
[0040] Furthermore, the supplementary response model analysis module includes a speech similarity set generation unit and a supplementary response model establishment unit;
[0041] The speech similarity set generation unit is used to extract the speech content in the control interaction event corresponding to each modal connection relationship as the control speech content, calculate the similarity between the control speech content and the target speech content in the target interaction event corresponding to the same type of response content, traverse and search all control interaction events recorded in the same modal connection relationship, calculate the similarity respectively, and generate the speech similarity set corresponding to the modal connection relationship;
[0042] The supplementary response model building unit is used to select the maximum and minimum values in the speech similarity set, and to build a supplementary response model for each type of response content recording the modal connection relationship corresponding to the speech features.
[0043] Furthermore, the prompt response content warning module includes a real-time output value analysis unit, a behavior similarity analysis unit, an evaluation coefficient calculation unit, and a prompt response content output unit;
[0044] The real-time output value analysis unit is used to calculate the similarity between the real-time speech content and the target speech content recorded by all models to obtain the similarity, and replace the minimum value in the supplementary response model to calculate the real-time output value;
[0045] The behavior similarity analysis unit is used to determine the magnitude relationship between the real-time output value and the output value of the supplementary response model;
[0046] The evaluation coefficient calculation unit is used to calculate the evaluation coefficient of the modal connection relationship corresponding to each type of effective response content;
[0047] The prompt response content output unit selects the valid response content corresponding to the minimum value of the evaluation coefficient calculated by the modal connection relationship as the prompt response content.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] 1. The present invention analyzes interactive operation events with irregular interaction modes based on historical records and selects voice content in valid interaction modes that can be used as a basis for judgment, thereby improving the accuracy of voice data acquisition and analysis during real-time interaction;
[0050] 2. In addition, corresponding evaluation models are generated for different interaction modes, which can effectively classify and judge when users are unable to touch the smart display in time under various interaction modes, and further evaluate the real-time interaction tendency for voice content in combination with various sensor data of the smart display, and finally achieve effective reminders when the touch response is not timely or the response is unsuccessful, thereby improving the fluency of the smart display under modal interaction and improving the user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 The present invention is a structural schematic diagram of a multi-source data interactive control system for an intelligent display. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] Example: Figure 1 As shown, the present invention provides a multi-source data interactive control system and method for an intelligent display, and a multi-source data interactive control method for an intelligent display, the method comprising the following steps:
[0054] Step S100: Filtering the smart display's records of interactive operation events that are not single-modal. An interactive operation event refers to an event that uses voice data to wake up the smart display for interactive operation. The interactive operation event records the voice content, user operation content, and smart display response content. Non-single-modal refers to recording the same smart display response content but with different modal interaction modes in each corresponding interactive operation event.
[0055] Step S200: Marking the interactive operation event corresponding to each type of response content when the number of recorded modal interaction modes is 1 as a first interactive event, extracting the voice content in the first interactive event as the first voice content corresponding to the response content type; determining the target voice content based on the first voice content analysis; and marking the interactive operation event corresponding to the target voice content as the target interactive event;
[0056] Step S300: extracting the remaining interactive operation events corresponding to the same type of response content except the target interactive event as reference interactive events; and constructing a modal connection relationship of the reference interactive events;
[0057] Step S400: extracting the speech content recorded for the interaction event based on the modal connection relationship, and generating a supplementary response model corresponding to each modal connection relationship;
[0058] Step S500: Based on the supplementary response model, the real-time captured voice content is obtained. If the smart display fails to successfully trigger the response content within the preset time period, the real-time voice content is analyzed based on the supplementary response model and a prompt response content is made.
[0059] Determining the target speech content based on the first speech content analysis includes the following specific steps:
[0060] Bind each first interaction event recorded with the same type of response content to the corresponding first voice content; sort the first interaction events from small to large according to the duration of the first voice content; extract the duration difference L1 of the first voice content recorded between the first first interaction event and the last first interaction event in the sequence;
[0061] If the duration difference L1 is less than or equal to the duration difference threshold L0, then extract the time length T between the time when the smart display initially receives the voice content and the time when the response content is displayed for each first interaction event record in the sequence, and calculate the duration dispersion Q of the first interaction event, Q = {[∑(T-T0) 2 ] / M} 1 / 2 Where M represents the number of first interaction events in the sequence; T0 represents the average duration of the first interaction event records; if the duration dispersion Q is less than or equal to the dispersion threshold Q0, all first interaction events in the sequence are output as target interaction events; if the duration dispersion Q is greater than the dispersion threshold Q0, the first interaction event corresponding to T less than T0 is selected as the target interaction event;
[0062] If the duration difference L1 is greater than the duration difference threshold L0, the duration difference between each first interaction event and the first first interaction event is calculated in sequence, and the first interaction event whose duration difference is less than the duration difference threshold is selected as the target interaction event;
[0063] The voice content corresponding to the target interaction event is used as the target voice content.
[0064] Analyzing the target interaction event indicates that the voice has successfully awakened the smart display for interactive operation; ensuring that effective evaluation criteria are recorded in the operation event of successful awakening interaction.
[0065] When the modal interaction mode is unique, it means that only the voice interaction mode is recorded. When there are multiple voice wake-ups and the modal interaction mode is unique, it means that the voice interaction mode does not effectively implement the interactive operation, so this part of the interaction events should be eliminated.
[0066] Constructing a modal connection relationship for a control interaction event includes the following steps:
[0067] Extract the user operation content of the control interaction event. The user operation content refers to other control modes other than the user voice control mode, such as the touch mode.
[0068] A monitoring interval is generated with the moment when the response content in the control interaction event is displayed as the end moment and the moment when the user triggers the voice control mode as the start moment. All control modes within the monitoring interval are extracted, and the modal connection relationship of each control interaction event is generated by connecting the voice control mode as the initial mode of the modal connection relationship and the other adjacent control modes in chronological order.
[0069] When the modal connection relationship in the control interaction event only has the voice control mode, that is, the voice control is repeated multiple times, the modal connection relationship is an independent modal relationship.
[0070] Step S400 includes the following specific steps:
[0071] Step S410: Each type of response content contains at least one type of modal connection relationship. The speech content in the comparison interaction event corresponding to each modal connection relationship is extracted as the comparison speech content. The similarity f1 between the comparison speech content and the target speech content in the target interaction event corresponding to the same type of response content is calculated. All comparison interaction events recorded with the same modal connection relationship are traversed and the similarity is calculated for each event, generating a speech similarity set F corresponding to the modal connection relationship.
[0072] Step S420: Select the maximum value fmax and the minimum value fmin in the speech similarity set F, and construct a supplementary response model Y corresponding to the speech features of each type of response content recording modal connection relationship; Y = (fmax-fmin) / fmax.
[0073] The purpose of analyzing the supplementary response model is to reduce the inconvenience caused to users when the voice control modal response cannot be successful in a timely manner. By establishing a supplementary response model, the voice content analysis under the modal connection relationship corresponding to different response contents can be effectively realized, thereby realizing effective analysis of the further response required by the smart display.
[0074] When the modal connection relationship recorded by the same type of response content includes both independent modal relationships and modal connection relationships of multiple modal connections, the recorded voice content data may be different; therefore, supplementary response models are established for different modal connection relationships.
[0075] Step S500 includes the following specific steps:
[0076] Obtaining real-time speech content and substituting it into several supplementary response models stored in the intelligent display, extracting the real-time speech content that satisfies the modal connection relationship corresponding to the supplementary response model. Satisfying the supplementary response model means calculating the similarity between the real-time speech content and the target speech content recorded by all models, and replacing the minimum value fmin in the supplementary response model to obtain the real-time output value. When the real-time output value is less than or equal to the output value of the corresponding supplementary response model, it is considered satisfied.
[0077] When the real-time speech content satisfies a unique supplementary response model, extracting the response content corresponding to the modal connection relationship recorded by the supplementary response model as the prompt response content;
[0078] When the supplementary response model satisfied by the real-time voice content is not unique, the response content corresponding to the supplementary response model is extracted as the effective response content, and the target interaction event corresponding to the effective response content in the historical record is obtained. The user screen after the voice content of each type of effective response content corresponding to the target interaction event record is obtained. The user screen is obtained by the smart display based on the preset access permission of the monitoring device.
[0079] Extract the real-time user screen after the real-time voice content is sent and before responding to other control modes, and calculate the user behavior similarity Z with the target interaction event recorded by each valid response content. For example, in the AI fitness function of the smart display, the display turns on the camera to obtain the user's body language for feature extraction, and then performs feature quantification and standardization. Finally, the Euclidean distance can be used to calculate the similarity to realize data analysis; based on the user behavior similarity and the real-time output value Y0 of the supplementary response model in which the corresponding valid response content is located, calculate the evaluation coefficient p of the modal connection relationship corresponding to each type of valid response content, p = k1*(1 / Z)+k2*Y0; where k1 and k2 represent the corresponding reference coefficients; set by the system;
[0080] The effective response content corresponding to the minimum value of the modal connection relationship evaluation coefficient is selected as the prompt response content. The prompt response content refers to the prompt content that the smart display analyzes and judges when the user fails to effectively trigger a response in the voice control mode, thereby improving the user's experience of the smart device due to the increase in multimodal interaction.
[0081] A multi-source data interaction control system for an intelligent display, the system comprising an interactive operation event screening module, a first interactive event marking module, a target interactive event analysis module, a modal connection relationship construction module, a supplementary response model analysis module, and a prompt response content warning module;
[0082] The interactive operation event screening module is used to screen interactive operation events recorded by the smart display that are not executed in a single mode;
[0083] The first interaction event marking module is used to mark the interaction operation event when the number of recorded modal interaction modes corresponding to each type of response content is 1 as the first interaction event;
[0084] The target interaction event analysis module is used to mark the interaction operation event corresponding to the target voice content as the target interaction event;
[0085] The modal connection relationship construction module is used to construct the modal connection relationship of the control interaction event;
[0086] The supplementary response model analysis module is used to generate a supplementary response model corresponding to each modal connection relationship;
[0087] The prompt response content warning module is used to analyze the real-time voice content based on the supplementary response model and make prompt response content.
[0088] The target interaction event analysis module includes an interaction event sorting unit, a duration dispersion calculation unit, and a target interaction event determination unit;
[0089] The interaction event sorting unit is used to bind each first interaction event recorded with the same type of response content to the corresponding first voice content; and sort the first interaction events from small to large according to the time length of the first voice content;
[0090] The duration dispersion calculation unit is used to calculate the duration dispersion of the first interaction event;
[0091] The target interaction event determining unit is configured to determine a first interaction event that meets the requirements as a target interaction event.
[0092] The supplementary response model analysis module includes a speech similarity set generation unit and a supplementary response model establishment unit;
[0093] The speech similarity set generation unit is used to extract the speech content in the control interaction event corresponding to each modal connection relationship as the control speech content, calculate the similarity between the control speech content and the target speech content in the target interaction event corresponding to the same type of response content, traverse and search all control interaction events recorded in the same modal connection relationship, calculate the similarity respectively, and generate the speech similarity set corresponding to the modal connection relationship;
[0094] The supplementary response model building unit is used to select the maximum and minimum values in the speech similarity set, and to build a supplementary response model for each type of response content recording the modal connection relationship corresponding to the speech features.
[0095] The prompt response content warning module includes a real-time output value analysis unit, a behavior similarity analysis unit, an evaluation coefficient calculation unit and a prompt response content output unit;
[0096] The real-time output value analysis unit is used to calculate the similarity between the real-time speech content and the target speech content recorded by all models to obtain the similarity, and replace the minimum value in the supplementary response model to calculate the real-time output value;
[0097] The behavior similarity analysis unit is used to determine the magnitude relationship between the real-time output value and the output value of the supplementary response model;
[0098] The evaluation coefficient calculation unit is used to calculate the evaluation coefficient of the modal connection relationship corresponding to each type of effective response content;
[0099] The prompt response content output unit selects the valid response content corresponding to the minimum value of the evaluation coefficient calculated by the modal connection relationship as the prompt response content.
[0100] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A multi-source data interactive control method for an intelligent display, characterized by: The method comprises the following steps: Step S100: Filtering the smart display's records of interactive operation events that perform non-single modality, wherein the interactive operation event refers to an event that uses voice data to wake up the smart display for interactive operation, and the interactive operation event records the voice content, user operation content, and smart display response content. Non-single modality refers to recording different modal interaction modes in each interactive operation event when the smart display response content is the same; Step S200: Marking the interactive operation event corresponding to each type of response content when the number of recorded modal interaction modes is 1 as a first interactive event, extracting the voice content in the first interactive event as the first voice content corresponding to the response content type; determining the target voice content based on the first voice content analysis; and marking the interactive operation event corresponding to the target voice content as the target interactive event; Step S300: extracting the remaining interactive operation events corresponding to the same type of response content except the target interactive event as reference interactive events; and constructing a modal connection relationship of the reference interactive events; Step S400: extracting the speech content recorded for the interaction event based on the modal connection relationship, and generating a supplementary response model corresponding to each modal connection relationship; Step S500: Based on the supplementary response model, the real-time captured voice content is obtained. If the smart display fails to successfully trigger the response content within the preset time period, the real-time voice content is analyzed based on the supplementary response model and a prompt response content is made.
2. The multi-source data interactive control method for an intelligent display according to claim 1, characterized in that: Determining the target speech content based on the first speech content analysis includes the following specific steps: Data-binding each first interaction event recorded with the same type of response content with the corresponding first voice content; and sorting the first interaction events from small to large according to the duration of the first voice content; Extract the duration difference L1 of the first speech content recorded between the first interaction event and the last interaction event in the sequence; If the duration difference L1 is less than or equal to the duration difference threshold L0, then extract the time length T between the time when the smart display initially receives the voice content and the time when the response content is displayed for each first interaction event record in the sequence, and calculate the duration dispersion Q of the first interaction event, Q = {[∑(T-T0) 2 ] / M} 1 / 2 ; Where M represents the number of first interaction events in the sequence; T0 represents the average duration of the first interaction event records; if the duration dispersion Q is less than or equal to the dispersion threshold Q0, all first interaction events in the output sequence are target interaction events; If the duration dispersion Q is greater than the dispersion threshold Q0, the first interaction event corresponding to T less than T0 is selected as the target interaction event; If the duration difference L1 is greater than the duration difference threshold L0, the duration difference between each first interaction event and the first first interaction event is calculated in sequence, and the first interaction event whose duration difference is less than the duration difference threshold is selected as the target interaction event; The voice content corresponding to the target interaction event is used as the target voice content.
3. The multi-source data interactive control method for an intelligent display according to claim 1, characterized in that: The construction of the modal connection relationship of the control interaction event includes the following steps: Extracting user operation content of the control interaction event, where the user operation content refers to other control modes excluding the user voice control mode; A monitoring interval is generated with the moment when the response content in the control interaction event is displayed as the end moment and the moment when the user triggers the voice control mode as the start moment. All control modes within the monitoring interval are extracted, and the modal connection relationship of each control interaction event is generated by connecting the voice control mode as the initial mode of the modal connection relationship and the other adjacent control modes in chronological order.
4. The multi-source data interactive control method for an intelligent display according to claim 2, characterized in that: The step S400 includes the following specific steps: Step S410: Each type of response content contains at least one type of modal connection relationship. The speech content in the comparison interaction event corresponding to each modal connection relationship is extracted as the comparison speech content. The similarity f1 between the comparison speech content and the target speech content in the target interaction event corresponding to the same type of response content is calculated. All comparison interaction events recorded with the same modal connection relationship are traversed and the similarity is calculated for each event, generating a speech similarity set F corresponding to the modal connection relationship. Step S420: Select the maximum value fmax and the minimum value fmin in the speech similarity set F, and construct a supplementary response model Y corresponding to the speech features of each type of response content recording modal connection relationship; Y = (fmax-fmin) / fmax.
5. The multi-source data interactive control method for an intelligent display according to claim 4, characterized in that: The step S500 includes the following specific steps: Obtaining real-time voice content and substituting it into several supplementary response models stored in the intelligent display, extracting the real-time voice content that satisfies the modal connection relationship corresponding to the supplementary response model, wherein satisfying the supplementary response model means calculating the similarity between the real-time voice content and the target voice content recorded by all models, and replacing the minimum value fmin in the supplementary response model to obtain the real-time output value. When the real-time output value is less than or equal to the output value of the corresponding supplementary response model, it is considered satisfied; When the real-time speech content satisfies a unique supplementary response model, extracting the response content corresponding to the modal connection relationship recorded by the supplementary response model as the prompt response content; When the supplementary response model satisfied by the real-time voice content is not unique, the response content corresponding to the supplementary response model is extracted as the effective response content, and the target interaction event corresponding to the effective response content is recorded in the historical record. The user screen after the voice content of each type of effective response content corresponding to the target interaction event record is obtained. The user screen is obtained by the intelligent display based on the preset access rights of the monitoring device. Extract the real-time user screen after the real-time voice content is sent and before responding to other control modes, and calculate the user behavior similarity Z with the target interaction event recorded by each valid response content. Based on the user behavior similarity and the real-time output value Y0 of the supplementary response model in which the corresponding valid response content is located, calculate the evaluation coefficient p of the modal connection relationship corresponding to each type of valid response content, p = k1*(1 / Z)+k2*Y0; where k1 and k2 represent corresponding reference coefficients; The effective response content corresponding to the minimum value of the evaluation coefficient calculated by the modal connection relationship is selected as the prompt response content.
6. A multi-source data interactive control system for an intelligent display, comprising: The system includes an interactive operation event screening module, a first interactive event marking module, a target interactive event analysis module, a modal connection relationship construction module, a supplementary response model analysis module, and a prompt response content warning module; The interactive operation event screening module is used to screen interactive operation events recorded by the smart display that are not executed in a single mode; The first interaction event marking module is used to mark the interaction operation event when the number of recorded modal interaction modes corresponding to each type of response content is 1 as the first interaction event; The target interaction event analysis module is used to mark the interaction operation event corresponding to the target voice content as the target interaction event; The modal connection relationship construction module is used to construct the modal connection relationship of the control interaction event; The supplementary response model analysis module is used to generate a supplementary response model corresponding to each modal connection relationship; The prompt response content warning module is used to analyze the real-time voice content based on the supplementary response model and generate prompt response content.
7. The multi-source data interactive control system for an intelligent display according to claim 6, characterized in that: The target interaction event analysis module includes an interaction event sorting unit, a duration dispersion calculation unit, and a target interaction event determination unit; The interaction event sorting unit is used to data-bind each first interaction event recorded with the same type of response content with the corresponding first voice content; and sorting the first interaction events from small to large according to the duration of the first voice content; The duration dispersion calculation unit is used to calculate the duration dispersion of the first interaction event; The target interaction event determining unit is configured to determine a first interaction event that meets the requirements as a target interaction event.
8. The multi-source data interactive control system for an intelligent display according to claim 6, characterized in that: The supplementary response model analysis module includes a speech similarity set generation unit and a supplementary response model establishment unit; The speech similarity set generation unit is used to extract the speech content in the control interaction event corresponding to each modal connection relationship as the control speech content, calculate the similarity between the control speech content and the target speech content in the target interaction event corresponding to the same type of response content, traverse and search all control interaction events recorded in the same modal connection relationship, calculate the similarity respectively, and generate the speech similarity set corresponding to the modal connection relationship; The supplementary response model building unit is used to select the maximum value and the minimum value in the speech similarity set, and build a supplementary response model that records the speech features corresponding to the modal connection relationship of each type of response content.
9. The multi-source data interactive control system for an intelligent display according to claim 6, characterized in that: The prompt response content warning module includes a real-time output value analysis unit, a behavior similarity analysis unit, an evaluation coefficient calculation unit and a prompt response content output unit; The real-time output value analysis unit is used to calculate the similarity between the real-time speech content and the target speech content recorded by all models to obtain the similarity, and replace the minimum value in the supplementary response model to calculate the real-time output value; The behavior similarity analysis unit is used to determine the magnitude relationship between the real-time output value and the output value of the supplementary response model; The evaluation coefficient calculation unit is used to calculate the evaluation coefficient of the modal connection relationship corresponding to each type of valid response content; The prompt response content output unit selects the valid response content corresponding to the minimum value of the modal connection relationship calculation evaluation coefficient as the prompt response content.
Citation Information
Patent Citations
Semantic information description implementation method and device
CN116432657A
Natural language intelligent large model interaction system and method based on AI analysis
CN118246432A
Context-aware control for smart devices
WO2019195799A1
Voice interaction method and apparatus, device, medium, and vehicle
WO2024222046A1