Multimodal pet interactive panel direct key-in initiation method and system
By setting up a shortcut key on the touchscreen to launch the pet interaction panel, combined with real-time monitoring and simulation of frog communication, multimodal feedback is provided, solving the problems of complex operation and lack of personalization in traditional systems, and achieving a fast, accurate, and personalized interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional multimodal pet interaction systems are complex to operate, lack intuitive command input methods, have low recognition accuracy, cannot be personalized, and are difficult to adapt to new technologies and equipment, affecting user experience and scalability.
The pet interaction panel can be activated by directly typing preset shortcut keys on the touchscreen. The input is monitored and recognized in real time, simulating the communication process of frog calls. Personalized interaction is provided through a multimodal feedback mechanism, supporting multiple interaction methods such as touch, voice, and vision.
It enables fast and accurate command input, enriches and personalizes the interactive experience, enhances the immediacy, smoothness and immersion of the interaction, and has wide applicability and scalability.
Smart Images

Figure CN119917171B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of pet interaction, and particularly to a multi-modal pet interaction panel direct key input starting method and system. BACKGROUND
[0002] Traditional multi-modal pet interaction system designs have complex operation interfaces and processes. Users need to go through multiple steps when starting the system or switching interaction modes, which reduces the user's convenience. Especially for pet owners who are not familiar with electronic devices, this complexity may become an obstacle for them to use such systems.
[0003] In terms of instruction input, traditional systems lack intuitive and efficient input methods. Users may need to input instructions through tedious menu navigation, specific key combinations, or complex voice commands, which not only increases the operation difficulty, but also may cause input errors, affecting the smoothness and accuracy of interaction. Traditional systems may still have problems of low accuracy or slow response speed in recognizing user inputs such as voice, gestures, facial expressions, etc. This may cause the system to fail to accurately understand the user's intention or fail to provide timely feedback, thus affecting the user's interaction experience.
[0004] In addition, traditional multi-modal pet interaction systems lack sufficient personalization and customization options. Users may not be able to adjust the interaction content, interface style, or feedback method according to their own preferences or the characteristics of their pets, which makes the system's interaction experience relatively single and difficult to meet the needs of different users. Traditional systems may only support specific devices or platforms, which limits the user's device selection and use scenarios. At the same time, with the continuous development of technology, it may be difficult to adapt to new recognition technologies or interaction methods, resulting in deficiencies in scalability and sustainability. SUMMARY
[0005] The technical problem to be solved by the present application is to provide a multi-modal pet interaction panel direct key input starting method and system, which realizes fast and accurate instruction input and provides rich and personalized interaction experience through a multi-modal feedback mechanism.
[0006] To solve the above technical problems, the technical solution of the present application is as follows:
[0007] In a first aspect, a multi-modal pet interaction panel direct key input starting method is provided, the method comprising:
[0008] Pre-set a plurality of specific instructions as shortcut keys for starting the pet interaction panel;
[0009] Directly enter the set shortcut keys through the touch screen, and monitor and identify in real time whether the input shortcut keys match the pre-set shortcut keys to obtain a matching result;
[0010] According to the matching result, start the pet interaction panel and display the multimodal interaction interface containing multiple interaction modes;
[0011] After the multimodal interaction interface is displayed, the parameters of the initial frog sound search are set, including the search space and the number of iterations;
[0012] Through real-time monitoring and analysis of user interaction operations on the touch screen, including touch and voice input, and according to the user's interaction operation, iterative search is performed to simulate the frog's call communication process, constantly listening and learning the user's preferences to identify and determine the user's interaction mode to obtain the recognition result;
[0013] According to the recognition result, the content of the multimodal interaction interface is updated in real time to reflect the user's preferred interaction mode, and the interaction result is fed back to the user through visual, auditory or tactile sensory channels.
[0014] Further, the set shortcut keys are directly entered through the touch screen, and it is monitored and identified in real time whether the input shortcut keys match the preset shortcut keys to obtain a matching result, including:
[0015] Send instructions to the touch screen driver to activate continuous monitoring of user touch input, including the location of the touch point, the type of touch event, and the duration of the touch;
[0016] When the touch screen driver captures a touch event, the corresponding event data is received, and it is determined whether the touch point is located in the preset shortcut key input area. If the touch event occurs in the shortcut key input area, the character sequence input by the user is read through the touch screen driver;
[0017] Start the regular expression matching process to iterate through the preset shortcut key list and match the character sequence input by the user. If a match is found, perform the operation associated with the matching shortcut key, including starting the pet interaction panel. If no match is found after iterating through the list, the monitoring state is maintained, and the user is prompted to re-enter to obtain a matching result.
[0018] Further, according to the matching result, start the pet interaction panel and display the multimodal interaction interface containing multiple interaction modes, including:
[0019] According to the matching result, send a start instruction to the pet interaction panel system, containing the user's identity and interaction mode preference;
[0020] The pet interaction panel system receives the start instruction, extracts the resources of the multimodal interaction interface, including image, audio, video files and interaction logic script;
[0021] After the resource extraction is completed, the multimodal interaction interface is initialized, including setting the layout, color, font visual elements of the interface and configuring the audio output, touch input interaction mode;
[0022] After initializing the multi-modal interaction interface, the pet interaction panel system displays the multi-modal interaction interface to the user through the touch screen.
[0023] Further, the multi-modal interaction interface provides the user with multiple ways to interact with the pet through touch, voice, and vision.
[0024] Further, the touch screen is used to listen to and analyze the user's interaction operations in real time, including touch and voice input, and iterative search is performed according to the user's interaction operations, the croaking communication process of the frog is simulated, the user's preferences are constantly listened to and learned, and the user's interaction mode is identified and determined to obtain an identification result, including:
[0025] The real-time listening function and voice recognition function of the touch screen are started, and when a touch event or voice input is detected, the user's operation intention is analyzed, including clicking a button, sliding a slider, and speaking a specific instruction;
[0026] Iterative search is performed according to the detected touch event and voice input, the croaking communication process of the frog is simulated, and in the iterative process, the fitness of the current interaction mode is constantly evaluated to obtain a fitness evaluation result;
[0027] According to the fitness evaluation result, the user's preferred interaction mode is learned and identified, and when a preset number of iterations is reached, the identification result is output, including the user's preferred interaction mode and commonly used operation instructions.
[0028] Further, the calculation formula of the fitness of the current interaction mode is:
[0029]
[0030] Where F represents the fitness of the current interaction mode, a1, a2, a3, a4 represent weight coefficients, b1, b2, b3, b4, b5 represent adjustment coefficients, T r represents the reaction time, T m represents the maximum acceptable reaction time, d1, d2, d3, d4, d5 represent the offset, N i represents the number of interactions, N e represents the preset number of interactions, T u represents the usage time within the evaluation period, T h represents the usage time threshold, P p represents the proportion of positive feedback, N f represents the total number of feedback given within the evaluation period, N n represents the minimum number of feedback.
[0031] Further, according to the recognition result, the content of the multi-modal interaction interface is updated in real time to reflect the user's preferred interaction mode, and the interaction result is fed back to the user through visual, auditory or tactile sensory channels, including:
[0032] According to the recognition result, the content matched with the user's preference is generated, including changing the arrangement of buttons, displaying different images or animations, and updating the content matched with the user's preference to the multi-modal interaction interface in real time;
[0033] The content matched with the user's preference is displayed on the touch screen, including image and animation visual elements, to obtain the recognition result and the interface content;
[0034] According to the recognition result and the interface content, audio feedback corresponding to the user's interaction is generated, including cheers, encouragement, and pet sounds.
[0035] In the second aspect, the multi-modal pet interaction panel directly enters the start system, including: an input recognition module for receiving and recognizing the user's direct input content, including text, symbols, numbers, and user input captured through the touch screen and microphone, including gestures, speech or facial expressions, to obtain the recognition result;
[0036] A preference analysis module for analyzing the recognition result to determine the user's preference and intention, including the user's preferred interaction mode, pet type and interaction scene, to obtain the analysis result;
[0037] An interface update module for updating the content of the multi-modal interaction interface in real time according to the analysis result of the preference analysis module, including changing the arrangement of buttons, displaying images or animations matched with the user's preference, adjusting the color and layout of the interface;
[0038] A feedback generation module for generating corresponding audio, visual or tactile feedback according to the interface content, including cheers, encouragement, pet sounds, dynamic images or vibration feedback;
[0039] An output display module including a touch screen, a loudspeaker and a vibration motor for displaying the updated interface content to the user and playing or providing the generated feedback;
[0040] A control logic module for coordinating the work of each module, responding to the user's input, and providing feedback.
[0041] In the third aspect, a computing device includes:
[0042] One or more processors;
[0043] A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the method.
[0044] In a fourth aspect, a computer readable storage medium stores a program which, when executed by a processor, implements the method.
[0045] The above scheme of the present application at least has the following beneficial effects:
[0046] By pre-setting multiple specific instructions as start shortcuts, the user only needs to simply enter these shortcuts on the touch screen to quickly start the pet interaction panel, without going through a complex operation process or navigating menus, improving the convenience of use. The iterative search mechanism simulating the communication process of the frog's croaking sound can continuously learn and adapt to the user's preferences by real-time monitoring and analyzing the user's interaction operations (including touch and voice input), so as to interact with the user in a more natural and personalized way, enhancing the realism and interest of the interactive experience.
[0047] By real-time monitoring and identifying the user's input shortcuts and subsequent iterative search based on the user's interaction operations, the system can accurately identify the user's intentions and needs and respond quickly. This mechanism not only improves the accuracy of identification, but also shortens the response time of the system, making the interaction more smooth and immediate. According to the user's interaction operations and preferences, the system can update the content of the multi-modal interaction interface in real time to reflect the user's preferred interaction method. This personalized customization not only meets the needs of different users, but also increases the diversity and interest of the interaction.
[0048] The system provides feedback on the interaction results to the user through various sensory channels such as vision, hearing or touch, such as displaying dynamic images, playing sound effects or providing vibration feedback, etc. This rich sensory feedback mechanism allows the user to be more immersed in the interaction with the pet and enjoy a more realistic and lively interactive experience. The touch screen input and iterative search mechanism has wide applicability and scalability. With the continuous development of technology, the system can easily integrate new recognition technologies or interaction methods, while supporting multiple devices and platforms to meet the diverse needs of users. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 is a flowchart of the direct key-in starting method of the multi-modal pet interaction panel provided by the embodiment of the present application.
[0050] Figure 2 is a schematic diagram of the direct key-in starting system of the multi-modal pet interaction panel provided by the embodiment of the present application. DETAILED DESCRIPTION
[0051] Exemplary embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0052] As shown in Figure 1 The embodiments of the present application propose a multimodal pet interaction panel direct keying start method, which comprises the following steps:
[0053] Step 11, a plurality of specific instructions are pre-set as shortcut keys for starting the pet interaction panel;
[0054] Step 12, the set shortcut keys are directly keyed in through the touch screen, and the input shortcut keys are monitored and identified in real time to determine whether they match the pre-set shortcut keys, and a matching result is obtained;
[0055] Step 13, according to the matching result, the pet interaction panel is started, and a multimodal interaction interface containing multiple interaction modes is displayed;
[0056] Step 14, after the multimodal interaction interface is displayed, the parameters of the initial frog sound search are set, including the search space and the number of iterations;
[0057] Step 15, the touch screen listens to and parses the user's interaction operations in real time, including touch and voice input, and iteratively searches according to the user's interaction operations, simulates the frog's call communication process, constantly listens to and learns the user's preferences, identifies and determines the user's interaction mode, and obtains an identification result;
[0058] Step 16, according to the identification result, the content of the multimodal interaction interface is updated in real time to reflect the user's preferred interaction mode, and the interaction result is fed back to the user through the visual, auditory or tactile sensory channels.
[0059] In the embodiments of the present application, by pre-setting the shortcut keys, the user can quickly and directly start the pet interaction panel without going through a cumbersome start process or searching for a start button, improving the convenience and efficiency of use. This design enables the pet owner to quickly interact with the pet at any time, enhancing the immediacy of the interaction experience.
[0060] Step 12, the direct keying method of the touch screen is intuitive and easy to operate, and the user only needs to touch the screen to input the shortcut keys. The real-time monitoring and identification mechanism ensures that the system can quickly respond to the user's input and accurately determine whether it matches the pre-set shortcut keys, thereby quickly starting the pet interaction panel, reducing the user's waiting time and improving the system's response speed.
[0061] Step 13, according to the matching result, quickly start the pet interaction panel and show a variety of interactive modes, providing users with rich choices. The design of the multi-modal interaction interface allows users to interact with the pet through visual, auditory, and other sensory channels, enhancing the diversity and immersion of the interactive experience.
[0062] Step 14, by setting the parameters of the initial frog sound search, the system can conduct targeted searches to more efficiently identify user preferences and interaction methods. Reasonable search space and iteration number settings ensure the accuracy of the search while avoiding unnecessary waste of computing resources, improving the overall performance of the system.
[0063] Step 15, real-time monitoring and analysis of user interaction operations allows the system to immediately obtain user feedback and needs. Through iterative search and simulation of the process of frog sound communication, the system can continuously learn and adapt to user preferences, thereby more accurately identifying and determining user interaction methods. This mechanism not only improves the accuracy of identification but also enhances the sensitivity and adaptability of the system to user needs.
[0064] Step 16, real-time updating of the content of the multi-modal interaction interface according to the identification results makes the interface more in line with user preferences and needs. Feedback of interaction results through visual, auditory, or tactile sensory channels not only enhances the richness and realism of the interactive experience but also allows users to more intuitively feel the effects of interaction with the pet, thereby further improving user satisfaction and loyalty.
[0065] In a preferred embodiment of the present application, the above-mentioned step 11, pre-setting a plurality of specific instructions as shortcut keys for starting the pet interaction panel, can include:
[0066] In an embodiment of the present application, a set of specific instructions to be used as shortcut keys for starting the pet interaction panel is determined. These instructions can be a combination of keys on the keyboard (such as Ctrl+Alt+P) or a specific button combination on the handle (such as LB+RB). The UI interface of the pet interaction panel is developed, including the image and status display of the pet, interaction options, etc. The display logic and interaction logic of the panel are determined, such as how to respond to player input and how to update the pet's status. The pet interaction panel is integrated into the game system or application supported by the gamepad to ensure that the panel can be displayed and interacted normally.
[0067] The gamepad firmware includes a monitoring mechanism to capture input signals from the controller. For keyboard shortcuts, monitoring needs to be implemented at the application level. When an input signal is detected, it is matched against a predefined set of shortcut commands. If a match is successful, the pet interaction panel's startup logic is triggered. Once a valid shortcut command is recognized, the pet interaction panel is launched immediately. Ensure the panel's display and interaction logic function correctly, such as loading the pet's image, displaying status information, and responding to player actions. The gamepad's settings menu provides a configuration interface for shortcut commands. This allows players to customize or modify shortcut commands to suit different gaming habits and needs. Player-configured shortcut commands are saved to the controller's memory or the application's configuration file. Ensure the configuration information is restored when the controller restarts or the application reopens. Conflict detection is performed during shortcut command configuration to avoid conflicts between different shortcut commands. If a conflict is detected, the player is prompted to reconfigure or select an alternative command.
[0068] In a preferred embodiment of the present invention, step 12 above, which involves directly inputting a preset shortcut key via a touchscreen and monitoring and identifying in real time whether the input shortcut key matches a preset shortcut key to obtain a matching result, may include:
[0069] Step 121: Send a command to the touch screen driver to activate the continuous monitoring function for user touch input. The monitoring content includes the location of the touch point, the type of touch event, and the duration of the touch.
[0070] Step 122: When the touch screen driver captures a touch event, it receives the corresponding event data and determines whether the touch point is located within the preset shortcut key input area. If the touch event occurs in the shortcut key input area, the touch screen driver reads the character sequence input by the user.
[0071] Step 123: Initiate the regular expression matching process, traverse the preset shortcut key list and match it with the character sequence entered by the user. If a match is successful, execute the operation associated with the matched shortcut key, including launching the pet interaction panel; if no match is found after traversing the list, maintain the listening state and wait for the user to re-enter in order to obtain a matching result.
[0072] In the embodiment of the invention, the touch screen driver is loaded and ensured to communicate correctly with the hardware. The driver parameters such as touch sensitivity, sampling rate, etc. are configured. Instructions are sent to the touch screen driver through the operating system to enable continuous listening mode. The instructions contain specific information to be listened for, such as the location of the touch point (x, y coordinates), the type of touch event (e.g. press, swipe, lift), and the duration of the touch. A callback function is defined that will receive the touch event data captured by the touch screen driver. The signature of the callback function is ensured to match the one described in the API documentation so that the driver can call it correctly. The pointer or reference to the callback function is registered with the touch screen driver, ensuring successful registration and checking for any possible errors or status codes.
[0073] In the callback function, only necessary operations are performed, such as updating the user interface, logging touch data, or triggering specific events. Avoid performing time-consuming or complex tasks in the callback function. Move these tasks to asynchronous threads or background tasks for processing. If the callback function needs to handle a large amount of touch data, consider using efficient data structures for storage and processing, such as queues, hash tables, or arrays. Implement error handling logic in the callback function to handle possible exceptional situations, such as invalid touch event data or resource shortages. Log error messages or send error notifications for timely diagnosis and problem resolution. Test and debug the callback function thoroughly to ensure it works correctly in various situations and does not introduce any latency or loss issues.
[0074] Step 122, when the touch screen driver captures a touch event, the corresponding event data is received through the callback function. The data should contain information such as the location of the touch point, the type of event, and the timestamp. According to the location information of the touch point, it is judged whether the touch event occurs in the preset shortcut key input area. This involves comparing the coordinates of the touch point with the boundaries of the shortcut key input area. If the touch event occurs in the shortcut key input area, the character sequence input by the user is read through the touch screen driver. This involves converting the movement trajectory of the touch point into character input or recognizing specific touch gestures as input.
[0075] Step 123, load the preset shortcut key list, each shortcut key is represented as a specific character sequence or pattern. Start the regular expression matching process to match the user input character sequence with each shortcut key in the shortcut key list. Compare the user input character sequence with each shortcut key in the shortcut key list in turn. If a match is found, exit the loop; if no match is found after traversing the list, remain in listening state. If a match is found, perform the corresponding operation according to the matched shortcut key. For the shortcut key that starts the pet interaction panel, trigger the display and interaction logic of the panel. If the match fails, wait for the user to re-enter and repeat the matching process.
[0076] Assuming the user has preset a shortcut "P+I+N" (abbreviation of "Pet Interaction Now") for launching the pet interaction panel. The user touches the shortcut input area on the screen and inputs "P", "I", "N" in sequence. The touch screen driver captures these touch events and reads the character sequence "P+I+N" input by the user. The regular expression matching process is started to match "P+I+N" with the preset shortcut list. After finding the match, the system performs the operation associated with the "P+I+N" shortcut, i.e., launching the pet interaction panel.
[0077] The user can launch the pet interaction panel by directly touching the screen to input the shortcut without additional physical buttons or complex operation processes. The function of real-time monitoring and recognizing the input shortcut enables the user to get immediate feedback, enhancing the smoothness and intuitiveness of the interaction. The user can customize the shortcut according to their preferences and habits, improving the flexibility and personalization of the system. The system can continuously listen and recognize the user's input shortcut, allowing the user to launch the pet interaction panel at any time without interrupting the current game or application. Through the regular expression matching process, the system can efficiently process the user's input shortcut and quickly execute the associated operation. This reduces the response time and processing delay of the system, improving the overall running efficiency.
[0078] In a preferred embodiment of the present application, step 13, according to the matching result, launching the pet interaction panel and displaying the multi-modal interaction interface containing multiple interaction modes, can include:
[0079] Step 131, according to the matching result, sends a launch instruction to the pet interaction panel system, containing the user's identity and interaction mode preference;
[0080] Step 132, the pet interaction panel system receives the launch instruction, extracts the resources of the multi-modal interaction interface, including image, audio, video files and interaction logic scripts;
[0081] Step 133, after the resource extraction is completed, initialize the multi-modal interaction interface, including setting the layout, color, font visual elements of the interface and configuring the audio output, touch input interaction mode;
[0082] Step 134, after initializing the multi-modal interaction interface, the pet interaction panel system displays the multi-modal interaction interface to the user through the touch screen.
[0083] In the embodiment of the present application, the user's identity information and interaction mode preference are obtained from the user identification system or previous interaction process. The matching result includes the user's ID, name, pet type preference, and interaction mode preference (such as game, education, feeding, etc.). According to the matching result, a start instruction data packet containing the user's identity and interaction mode preference is constructed. The data packet includes the user's ID, interaction mode identifier, and other necessary parameters (such as pet type, interaction level, etc.). The start instruction data packet is sent to the pet interaction panel system through a network communication protocol (such as TCP / IP, HTTP, etc.).
[0084] Step 132, the pet interaction panel system listens to the start instruction from the user identification system. When receiving the start instruction, the data packet is parsed to obtain the user's identity and interaction mode preference. According to the user's identity and interaction mode preference, the required multi-modal interaction interface resources are located in the resource library. The resources may include image files (such as background pictures, pet icons), audio files (such as background music, pet sounds), video files (such as animations, tutorial videos), and interaction logic scripts.
[0085] Step 133, according to the interaction mode preference and the preset layout template, the overall layout of the interface is set, including the position and size of each element. The color scheme, font style and size of the interface are set to ensure the beauty and readability of the interface. Image resources such as background pictures and pet icons are loaded and displayed. Audio output devices such as speakers or earphones are initialized. Audio resources such as background music and pet sounds are loaded and prepared for playback. According to the interaction logic script, the input interaction mode of the touch screen is configured, such as touch point recognition and sliding operation response.
[0086] Step 134, the initialized multi-modal interaction interface is rendered to the touch screen. Ensure that the rendering of the interface is smooth without stuttering or flickering. Real-time listen to the user's touch input and process the input events according to the interaction logic script. Update the interface state to reflect the user's operation results, such as moving pet icons and playing animations. Feedback is provided to the user through audio, visual or tactile means, such as playing sound effects and displaying prompt information.
[0087] Assuming user A enjoys interactive games with virtual pets, previous matching results indicate that user A prefers "dog" as the pet type and likes the "catching frisbee" interaction mode. An initiation instruction is sent to the pet interaction panel, containing user A's ID and the identification of the "catching frisbee" mode. After receiving the instruction, the pet interaction panel system extracts the "dog" image, "catching frisbee" background image, dog barking sound, and interactive logic script of the frisbee game from the resource library. The interface is initialized, setting a green meadow as the background, displaying a cute dog icon, and configuring the speaker to play dog barking sounds and background music. At the same time, the touch screen's sliding operation is set to control the dog to move and catch the frisbee. After the interface is rendered, user A sees the dog and frisbee on the touch screen and controls the dog to move and catch the frisbee by sliding the screen. Each time the frisbee is caught, the system will play cheerful music and dog barking as feedback.
[0088] Through the multi-modal interactive interface, users can more intuitively interact with pets and enjoy a more immersive experience. Touch input and real-time feedback mechanisms enable users to interact with pets more naturally, increasing the fun and realism of the interaction. According to the user's identity and preferences, the system can automatically adjust the interactive interface and resources to meet the user's individual needs. By configuring and loading different resources, the system can easily support multiple interaction modes and pet types, improving the flexibility and scalability of the system.
[0089] In another preferred embodiment of the present application, the multi-modal interactive interface provides users with multiple ways to interact with pets through touch, voice, and vision, which can include:
[0090] In the embodiment of the present application, the pet interaction system is started, loading necessary system components and services. Load the images, audio, video files and interactive logic scripts required by the multi-modal interactive interface from the resource library. These resources include different expressions and action images of pets, pet sounds, background music, and interactive game animation videos and logic processing scripts. Identify the user's identity through face recognition, fingerprint recognition or other identity verification methods. According to the user's historical data or preset configuration, obtain the user's interaction mode preferences, such as preferred pet types and interactive games.
[0091] According to the user's preferences and preset layout templates, set the overall layout of the multi-modal interactive interface, including the positions and sizes of the pet display area, interactive game area, control buttons, etc. Set the color scheme, font style and size of the interface, and load and display the initial image and background image of the pet and other visual elements. Initialize the audio output device, such as the speaker or earphone, to prepare to play pet sounds, background music and other audio resources. Configure speech recognition and speech synthesis services so that users can interact with pets through voice.
[0092] The user interacts with the pet through the touch screen, such as clicking on the pet image to trigger the pet's reaction, dragging the pet to move, or sliding in the interactive game area, etc. The system updates the interface state in real time according to the user's touch input, such as changing the pet's expression, action, or updating the game progress, etc., and enhances the interactive experience through tactile feedback (such as vibration). The user speaks instructions or has a conversation with the pet through the microphone, and the system uses speech recognition technology to convert the user's voice into text. According to the recognized text content, the user's intention is judged, and the corresponding interactive logic is called to process. Using speech synthesis technology, the pet's response or the system's prompt information is converted into voice output, which is played to the user through the speaker or earphone. According to the interactive logic, the system updates the pet's image and animation in real time, such as the pet's expression change, action execution, etc., to provide visual interactive experience. During the interaction, the system can add visual effects such as flashing, gradient, zooming, etc. to enhance the attractiveness and interactive effect of the interface. Determine whether the interaction time meets the pre-set end condition. Detect whether the user has performed an end interaction operation, such as clicking the exit button or speaking an end instruction, etc. Release the resources occupied by the multi-modal interactive interface, such as image, audio, video files and memory, etc. Close the multi-modal interactive interface and return to the main interface of the system.
[0093] In a preferred embodiment of the present application, after the multi-modal interactive interface is displayed in step 14, the parameters of the initial frog sound search are set, including the search space and the number of iterations, which can include:
[0094] In an embodiment of the present application, the search space can involve the following aspects:
[0095] Visual parameters: such as the color range, brightness, contrast, etc. of the interface.
[0096] Audio parameters: such as the volume, pitch of background music, and variations of pet sounds, etc.
[0097] Interactive logic parameters: such as the sensitivity of the pet's reaction, game difficulty, switching threshold of interactive mode, etc.
[0098] For each parameter to be optimized, its minimum and maximum values are determined, thereby defining a continuous search space. The number of iterations determines how many times the frog sound search algorithm will run iterations to find the final solution.
[0099] In a preferred embodiment of the present application, step 15, the user's interactive operation is listened to and parsed in real time through the touch screen, including touch and voice input, and the iterative search is performed according to the user's interactive operation, simulating the process of frog sound communication, constantly listening to and learning the user's preferences to identify and determine the user's interactive mode to obtain the recognition result, which can include:
[0100] Step 151: Activate the real-time monitoring function and voice recognition function of the touch screen. When a touch event or voice input is detected, analyze the user's operation intention, including clicking a button, sliding a slider, or speaking a specific command.
[0101] Step 152: Perform iterative search based on detected touch events and voice input, simulating the communication process of frog calls. During the iteration process, continuously evaluate the fitness of the current interaction mode to obtain the fitness evaluation result.
[0102] Step 153: Based on the fitness assessment results, learn and identify the user's preferred interaction methods. When the preset number of iterations is reached, output the identification results, including the user's preferred interaction methods and commonly used operation commands.
[0103] In this embodiment of the invention, a touchscreen monitoring service and a speech recognition service are initialized. The touchscreen monitoring service is responsible for capturing user touch events, such as clicks and swipes; the speech recognition service is responsible for converting user speech input into text. Parameters such as the sensitivity and response speed of the touchscreen monitoring service, and the language model and recognition accuracy of the speech recognition service are configured to ensure accurate capture and parsing of user interactions. When the touchscreen detects a touch event or the speech recognition service recognizes speech input, an event handling function is triggered. This function is responsible for parsing the user's intent, such as determining which button was clicked, which slider was slid, or recognizing a specific command spoken by the user.
[0104] Step 152: Define a search space containing multiple possible interaction modes, such as different button layouts, slider ranges, and instruction sets. These interaction modes can be viewed as different variations of "frog croaking." Set initial parameters for the iterative search, such as the number of iterations and the fitness evaluation function. The fitness evaluation function is used to assess the degree of match between the current interaction mode and user preferences. In each iteration, select an interaction mode from the search space as the current mode. Based on the current mode, simulate the frog croaking communication process, i.e., display interface elements, respond to touch events, and respond to voice input according to the mode. Collect user interaction data in this mode, such as operation frequency, dwell time, and satisfaction feedback. Use the fitness evaluation function to evaluate the current mode and obtain the fitness evaluation result. Based on the fitness evaluation result, update the search space, eliminate modes with low fitness, and optimize modes with high fitness.
[0105] Step 153, during the iterative search process, the user's interaction data is continuously accumulated and analyzed and modeled using machine learning algorithms to learn the user's preferences. When the preset number of iterations is reached or the fitness evaluation result reaches a certain threshold, the iterative search is stopped. At this time, the system has learned the user's preferences and can identify the user's preferred interaction methods and commonly used operation instructions. The identification results are displayed to the user in a visual manner, such as displaying the user's preferred button layout, slider range, etc. At the same time, the identification results can also be used to optimize the interactive design of the system and improve the user experience.
[0106] Suppose a user is using a multi-modal interactive interface and often triggers a certain function by touching a certain area on the screen and often speaks a certain specific voice instruction to perform a certain operation. During the iterative search process, the system will gradually learn these preferences and preferentially display the user's preferred button layout and respond to the user's preferred voice instructions during interaction. For example, the user's commonly used function buttons will be placed in a more eye-catching position on the screen, and the system will respond quickly when the user speaks a specific instruction.
[0107] By real-time monitoring and analyzing the user's interaction operations, the system can more accurately understand the user's needs and preferences, thereby providing a more personalized and efficient interactive experience. Through iterative search and learning of user preferences, the system can continuously adapt to changes and needs of the user, keeping the interaction between the system and the user smooth and unobstructed. By identifying the user's preferred interaction methods and operation instructions, the system can more reasonably allocate resources, such as preferentially loading the user's commonly used function modules, reducing unnecessary resource consumption.
[0108] In a preferred embodiment of the present application, the formula for calculating the fitness of the current interaction mode is:
[0109]
[0110] Where F represents the fitness of the current interaction mode; α1, α2, α3, α4 represent weight coefficients; β1, β2, β3, β4, β5 represent adjustment coefficients; T r represents the reaction time; T m represents the maximum acceptable reaction time; δ1, δ2, δ3, δ4, δ5 represent the offset; N i represents the number of interactions; N e represents the preset number of interactions; T u represents the usage time within the evaluation period; T h represents the usage time threshold; P p represents the proportion of positive feedback; N f represents the total number of feedback given within the evaluation period; N n represents the minimum number of feedback.
[0111] In the embodiments of the present application, the weight coefficients (α1, α2, α3, α4), adjustment coefficients (β1, β2, β3, β4, β5), offset amounts (δ1, δ2, δ3, δ4, δ5), maximum acceptable reaction time (T m ), preset interaction number (N e ), usage time threshold (T h ), and minimum feedback number (N n ) and other parameters are read from the configuration file. The loaded parameters are checked to ensure that they are within a reasonable range, avoiding calculation errors or abnormalities. The reaction time (T r ), interaction number (N i ), usage time (T u ) in the evaluation period, positive feedback ratio (P p ), and total number of feedback given in the evaluation period (N f ) and other variables are initialized for subsequent data collection and calculation.
[0112] The touch screen and voice recognition module are used to monitor the user's interactive operation in real time, including touch events and voice input. The reaction time (T r ) of each interaction is recorded, and the interaction number (N i ) is accumulated. At the same time, the usage time (T u ) of the user in the evaluation period and the total number of feedback given (N f ), as well as the number of positive feedback, are recorded for calculating the positive feedback ratio (P p ).
[0113] Calculate the reaction time component
[0114] Calculate the interaction number component
[0115] Calculate the usage time component
[0116] Calculate the feedback component
[0117] Add all components to get the fitness F of the current interaction mode.
[0118] By calculating the fitness in real time and adjusting the interaction strategy, the system can more accurately meet the user's needs and preferences, providing a more personalized and smooth interaction experience. The fitness calculation formula takes into account multiple factors such as reaction time, interaction frequency, usage time, feedback, etc., allowing the system to more comprehensively and accurately evaluate the pros and cons of the interaction mode and make adjustments and optimizations accordingly, enhancing the system's adaptability. Through fitness evaluation, the system can identify the user's common interaction methods and operation instructions, and prioritize resource allocation such as loading speed and processing priority, improving the system's running efficiency and response speed.
[0119] In a preferred embodiment of the present application, the above-mentioned step 16, according to the recognition result, real-time updates the content of the multi-modal interaction interface to reflect the user's preferred interaction method, and feeds back the interaction result to the user through visual, auditory or tactile sensory channels, which can include:
[0120] Step 161, according to the recognition result, generates content that matches the user's preferences, including changing the arrangement of buttons, displaying different images or animations, and updating the content that matches the user's preferences to the multi-modal interaction interface in real time;
[0121] Step 162, through the touch screen, display the real-time updated content that matches the user's preferences, including image and animation visual elements, to obtain the recognition result and interface content;
[0122] Step 163, according to the recognition result and interface content, generate audio feedback corresponding to the user's interaction, including cheers, encouragement, and pet sounds.
[0123] In the embodiment of the present application, the latest user preference data is obtained from the user behavior analysis module, including common interaction methods, preferred content types (such as images, animations), favorite colors, layouts, etc. According to the user preference data, search for matching content in the content library, such as specific button arrangement, image or animation materials. Use the graphics processing engine to generate interface elements that match the user's preferences based on the matched content, such as adjusting the size, color, and arrangement of buttons, or loading the user's preferred images and animations.
[0124] The generated interface elements are updated to the corresponding positions of the multi-modal interaction interface in real time, ensuring that the user can see the content that meets their preferences every time they interact.
[0125] Step 162, render the updated interface elements into the display buffer of the touch screen, ensuring smooth display of images and animations. Listen to touch events of the touch screen, when the user touches the screen, according to the touch position and content layout, determine the user's intention and trigger the corresponding interaction logic. Through the touch screen, real-time updated interface content is displayed, including user triggered animation effects, image changes, etc., so that users can intuitively see the interaction results.
[0126] Step 163, prepare a variety of audio materials in advance, such as cheers, encouragement, pet sounds, etc., corresponding to different interaction results and situations. According to the recognition result and interface content, determine the user's current interaction situation and result, and select the matching audio material from the audio library. Use the audio playback engine to play the selected audio material, and deliver audio feedback to the user through the speaker or earphone, enhancing the immersion and interest of the interaction.
[0127] Assume that the user's preference is to like pet-themed interactive interfaces, and prefer to hear pet sounds when clicking buttons. After recognizing the user's preference, load pet-themed images and animations from the content library, such as dog icons, bone buttons, etc., and adjust the button arrangement to the user's usual way. When the user clicks the bone button on the touch screen, the dog eating the bone animation will be displayed on the interface, and the button arrangement will be adjusted according to the user's habit. Play the audio feedback of the dog's bark to make the user feel the fun of interacting with pets.
[0128] By real-time updating of content matching user preferences and providing visual, auditory and other multi-sensory feedback, users can experience a more personalized and immersive experience during interaction. Interactive interfaces and feedback mechanisms that meet user preferences can attract users to use the system more frequently, improving user satisfaction and loyalty. Adjusting the interface layout and button arrangement according to user preferences allows users to quickly find and trigger the desired interactive functions, improving interaction efficiency. Through audio feedback such as pet sounds, emotional communication between the system and the user is enhanced, making users more emotionally dependent and identified with the system. The system can be adjusted and optimized in real time according to user preferences, providing a basis and support for future personalized customization services.
[0129] As shown in Figure 2 The embodiments of the present application also provide a multi-modal pet interaction panel for directly entering to start the system, which comprises:
[0130] An input recognition module is configured to receive and recognize the direct input content of the user, including text, symbols, numbers, and user input captured through the touch screen and microphone, including gestures, voice or facial expressions, to obtain a recognition result.
[0131] a preference analysis module configured to analyze the recognition result and determine the user's preference and intention, including the user's preferred interaction mode, pet type and interaction scene, to obtain an analysis result;
[0132] an interface updating module configured to update the content of the multi-modal interaction interface in real time according to the analysis result of the preference analysis module, including changing the arrangement of buttons, displaying images or animations matching the user's preference, and adjusting the color and layout of the interface;
[0133] a feedback generation module configured to generate corresponding audio, visual or tactile feedback according to the interface content, including cheers, encouragement, pet sounds, dynamic images or vibration feedback;
[0134] an output display module including a touch screen, a speaker and a vibration motor, configured to display the updated interface content to the user and play or provide the generated feedback;
[0135] a control logic module configured to coordinate the work of the modules, respond to the user's input and provide feedback
[0136] It should be noted that the system is a system corresponding to the above method, and all implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0137] Embodiments of the present application also provide a computing device, including a processor and a memory storing a computer program, wherein the computer program is executed by the processor to perform the method described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0138] Embodiments of the present application also provide a computer readable storage medium storing instructions, which, when executed on a computer, cause the computer to perform the method described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0139] The above is the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should also be considered within the scope of protection of the present application.
Claims
1. A multi-modal pet interactive panel direct key-in launch method, characterized in that, The method comprises: S11, presetting a plurality of specific instructions as shortcut keys for starting the pet interaction panel; S12, directly entering the set shortcut keys through the touch screen, and monitoring and identifying in real time whether the input shortcut keys match the preset shortcut keys to obtain a matching result; S13, starting the pet interaction panel according to the matching result, and displaying a multimodal interaction interface containing a plurality of interaction modes; S14, after the multimodal interaction interface is displayed, setting initial parameters, including a search space and an iteration number, wherein the search space includes visual parameters and audio parameters; S15, listening to and analyzing the user's interaction operation in real time through the touch screen, including touch and voice input, and performing iterative search according to the user's interaction operation, constantly listening to and learning the user's preferences in the iterative search process, to identify and determine the user's interaction mode to obtain an identification result; including: S151, starting the real-time listening function and voice recognition function of the touch screen, analyzing the user's operation intention when detecting a touch event or voice input, including clicking a button, sliding a slider, and speaking a specific instruction; S152, performing iterative search according to the detected touch event and voice input, constantly evaluating the fitness of the current interaction mode in the iteration process to obtain a fitness evaluation result; S153, learning and identifying the user's preferred interaction mode according to the fitness evaluation result, and outputting the identification result when the preset iteration number is reached, including the user's preferred interaction mode and commonly used operation instructions; In step 152, a search space containing a plurality of interaction modes is defined; in each iteration, an interaction mode is selected from the search space as the current mode, interface elements are displayed according to the mode, touch events and voice inputs are responded to, user interaction data in the mode is collected, a fitness evaluation function is used to evaluate the current mode to obtain a fitness evaluation result, and the search space is updated according to the fitness evaluation result, eliminating low-fitness modes and optimizing high-fitness modes; In step 153, in the iterative search process, the user's interaction data is constantly accumulated, and machine learning algorithms are used to analyze and model these data to learn the user's preferences; when the preset iteration number is reached or the fitness evaluation result reaches a certain threshold, the iterative search is stopped, at this time, the system has learned the user's preferences and can identify the user's preferred interaction mode and commonly used operation instructions, and the identification result is displayed to the user in a visual manner. S16, updating the content of the multi-modal interactive interface in real time according to the recognition result to reflect the interactive mode preferred by the user, and feeding back the interactive result to the user through the visual, auditory or tactile sensory channel, including: generating content matched with the user's preference according to the recognition result, including changing the arrangement of buttons, displaying different images or animations, and updating the content matched with the user's preference to the multi-modal interactive interface in real time; displaying the content matched with the user's preference updated in real time through the touch screen, including image and animation visual elements, to obtain the recognition result and the interface content; generating audio feedback corresponding to the user interaction according to the recognition result and the interface content, including cheers, encouragement and pet sounds.
2. The multi-modal pet interactive panel direct key-in launch method of claim 1, wherein, The set shortcut keys are directly entered through the touch screen, and it is monitored and recognized in real time whether the input shortcut keys match the preset shortcut keys to obtain a matching result, including: sending an instruction to the touch screen driver to activate the continuous monitoring function of the user touch input, and the monitoring content includes the position of the touch point, the type of the touch event and the duration of the touch; when the touch screen driver captures the touch event, the corresponding event data is received, and it is judged whether the touch point is located in the preset shortcut key input area, if the touch event occurs in the shortcut key input area, the character sequence input by the user is read through the touch screen driver; starting a regular expression matching process to match the preset shortcut key list with the character sequence input by the user, if the matching is successful, the operation associated with the matching shortcut key is executed, including starting the pet interactive panel; if no matching result is obtained after traversing the list, the monitoring state is maintained, and the user is waited to re-input to obtain a matching result.
3. The multi-modal pet interactive panel direct key-in launch method of claim 2, wherein, According to the matching result, the pet interactive panel is started, and a multi-modal interactive interface containing multiple interactive modes is displayed, including: According to the matching result, an activation instruction is sent to the pet interactive panel system, including the user's identity and the interactive mode preference; The pet interactive panel system receives the activation instruction, extracts the resources of the multi-modal interactive interface, including image, audio, video files and interactive logic scripts; After the resource extraction is completed, the multi-modal interactive interface is initialized, including setting the layout, color, font visual elements of the interface and configuring the audio output and touch input interaction mode; After initializing the multi-modal interactive interface, the pet interactive panel system displays the multi-modal interactive interface to the user through the touch screen.
4. The multi-modal pet interactive panel direct key-in launch method of claim 3, wherein, The multi-modal interactive interface provides multiple ways for the user to interact with the pet through touch, voice and vision.
5. The multi-modal pet interactive panel direct key-in launch method of claim 4, wherein, The formula for calculating the fitness of the current interactive mode is: ; wherein, represents the fitness of the current interaction mode; , , , represents a weight coefficient; , , , , represents an adjustment coefficient; represents a reaction time; represents a maximum acceptable reaction time; , , , , represents an offset; represents the number of interactions; represents a preset number of interactions; represents the usage time within the evaluation period; represents a usage time threshold; represents a positive feedback ratio; represents the total number of feedback given within the evaluation period; represents a minimum feedback quantity.
6. A multi-modal pet interactive panel direct key-in launch system, the system implementing the method of any one of claims 1 to 5, characterized in that, including: an input recognition module for receiving and recognizing the direct input content of the user, including text, symbols, numbers and user input captured through the touch screen and microphone, including gestures, voice or facial expressions, to obtain a recognition result; a preference analysis module for analyzing the recognition result to determine the user's preference and intention, including the user's preferred interactive mode, pet type and interactive scene, to obtain an analysis result; interface updating module, configured to update content of the multi-modal interactive interface in real time according to the analysis result of the preference analysis module, including changing arrangement of buttons, displaying images or animations matching the user preference, adjusting color and layout of the interface; a feedback generation module, configured to generate corresponding audio, visual or tactile feedback according to the interface content, including cheering, encouraging, pet sound, dynamic image or vibration feedback; an output display module, including a touch screen, a speaker and a vibration motor, configured to display the updated interface content to the user and play or provide the generated feedback; a control logic module, configured to coordinate work of the modules, respond to input of the user and provide feedback.
7. A computing device, comprising: comprise: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a program, and the program is executed by the processor to implement the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Method of gesture rapid start for touch screen mobile phone
CN102339151A
Small program front-end page website building design method and system
CN119045790A