Intelligent equipment information processing method and system, terminal and storage medium
By acquiring scene data from smart devices, determining scene type and interaction theme information, and generating AI interaction services strongly related to the current scene, the problem of AI assistants being unable to understand context in existing technologies is solved, thereby improving interaction effects and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-04-14
AI Technical Summary
The AI assistants of existing smart devices cannot understand the context of the content, resulting in the inability to provide interactive services that are highly relevant to the current scenario, thus affecting the interaction effect and user experience.
By acquiring images, audio, text, and metadata from smart devices, the system determines the scene type and interaction theme information, and generates AI interaction services that are highly relevant to the current scene.
It enables interactive services that are highly relevant to the current scenario, improves the interaction effect and user experience, and provides a personalized and seamless interactive experience.
Smart Images

Figure CN121865040A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence interaction technology, and in particular to a method, system, terminal and storage medium for information processing of intelligent devices. Background Technology
[0002] Currently, various smart devices (such as smart TVs, smart tablets, and smart home control screens) are being used more and more widely. With the development of artificial intelligence (AI) technology, AI assistants have also become one of the core functions for enhancing the user experience of smart devices.
[0003] In the existing technology, some smart devices are equipped with AI assistants, but existing AI assistants can usually only provide interactive feedback based on the content input by the user in real time, and cannot understand the context of the corresponding content. Therefore, they cannot provide interactive services that are strongly related to the current scene, which affects the interaction effect and user interaction experience.
[0004] Therefore, the relevant technologies still need to be improved and developed. Summary of the Invention
[0005] The main purpose of this application is to provide a method, system, terminal and storage medium for information processing of intelligent devices, which aims to solve the technical problem that AI assistants in related technologies can usually only provide interactive feedback based on the content input by the user in real time, but cannot understand the context of the corresponding content, and therefore cannot provide interactive services that are strongly related to the current scene, thus affecting the interaction effect and user interaction experience.
[0006] To achieve the above objectives, the first aspect of this application provides a method for processing information in a smart device, wherein the method includes: In response to an interaction trigger signal, scene data corresponding to the smart device is acquired, wherein the scene data includes at least one of image data, audio data, text data and metadata, and the metadata is used to indicate application information corresponding to the application running in the smart device. Based on the above scenario data, determine the scenario type and interaction theme information that match the above scenario data; Based on the above scenario types and interaction theme information, generate AI interaction services corresponding to the above smart devices.
[0007] Optionally, the above method further includes: In response to the detection of a global trigger operation, an interactive trigger signal is generated.
[0008] Optionally, the above-mentioned response to the interaction trigger signal to obtain scene data corresponding to the smart device includes: In response to the interaction trigger signal, the screen of the smart device is captured to obtain image data, audio data is obtained according to the audio stream being played on the smart device, text data corresponding to the display interface of the smart device is identified and obtained, and metadata is determined according to the foreground application running on the smart device. The aforementioned metadata includes the package name, activity name, and window component information of the aforementioned foreground application.
[0009] Optionally, the above-mentioned determination of scene type and interaction theme information matching the above-mentioned scene data includes: Send the above scenario data to the preset analysis platform; The analysis platform described above is used to determine the type of scenario that matches the scenario data described above. Based on the above scenario types, a target topic analysis model is determined from multiple preset topic analysis models, and interactive topic information matching the above scenario data is determined based on the target topic analysis model.
[0010] Optionally, the above scene type is one of several preset types; The aforementioned preset types include media playback scenarios, game scenarios, application interface scenarios, and system desktop scenarios; The aforementioned interactive theme information is used to characterize the interactive theme for the aforementioned scenario data under the aforementioned scenario type.
[0011] Optionally, the above method further includes: In response to the save command input by the target object, an AI interactive theme control based on a preset format is generated for the aforementioned AI interactive service, and the aforementioned AI interactive theme control is stored in the data space corresponding to the aforementioned target object.
[0012] Optionally, the above method further includes: In response to detecting a scenario type that matches the aforementioned AI interaction service, the aforementioned AI interaction theme control is displayed on the aforementioned smart device; In response to the interaction signal triggered by the target object based on the AI interaction theme control, the AI interaction service is executed.
[0013] A second aspect of this application provides an intelligent device information processing system, wherein the system includes: The data acquisition module is used to acquire scene data corresponding to the smart device in response to the interaction trigger signal. The scene data includes at least one of image data, audio data, text data and metadata. The metadata is used to indicate the application information corresponding to the application running in the smart device. The data processing module is used to determine the scene type and interaction theme information that match the scene data based on the above scene data. The service generation module is used to generate AI interaction services corresponding to the smart devices mentioned above, based on the scenario types and interaction theme information.
[0014] A third aspect of this application provides a terminal, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of any of the above-described intelligent device information processing methods.
[0015] A fourth aspect of this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-described intelligent device information processing methods.
[0016] As can be seen from the above, the present application provides a method for processing information of a smart device. Specifically, in response to an interaction trigger signal, scene data corresponding to the smart device is acquired. The scene data includes at least one of image data, audio data, text data, and metadata. The metadata is used to indicate application information corresponding to the application running in the smart device. Based on the scene data, a scene type and interaction theme information matching the scene data are determined. Based on the scene type and interaction theme information, an AI interaction service corresponding to the smart device is generated.
[0017] Thus, during the use of smart devices, in response to interaction trigger signals, scene data that is context-dependent on the current interaction content and can characterize the current scene is acquired. This determines the scene type and interaction theme information that match the scene data, and then generates AI interaction services in real time and dynamically based on the scene type and interaction theme information. In this way, the generated AI interaction service is generated based on the context information of the user-input interaction content and is strongly correlated with the current scene of the smart device. Therefore, the solution of this application can provide interactive services that are strongly correlated with the current scene, which is beneficial to improving the interaction effect and user interaction experience. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating an information processing method for intelligent devices provided in an embodiment of this application; Figure 2This is a schematic diagram of an AI interactive service system architecture provided in an embodiment of this application; Figure 3 This is a schematic flowchart illustrating a smart device information processing method provided in an embodiment of this application. Figure 4 This is a schematic diagram illustrating a specific process for multimodal content analysis and topic extraction provided in an embodiment of this application; Figure 5 This is a schematic diagram of an AI-themed interactive details interface provided in an embodiment of this application; Figure 6 This is a schematic diagram of an AI card library interface provided in an embodiment of this application; Figure 7 This is a schematic diagram of the constituent modules of an intelligent device information processing system provided in an embodiment of this application; Figure 8 This is a block diagram illustrating the internal structure of a terminal provided in an embodiment of this application. Detailed Implementation
[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.
[0021] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0022] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0023] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0024] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to classification." Similarly, the phrases "if determined" or "if classified to [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once classified to [the described condition or event]," or "in response to classification to [the described condition or event]."
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0026] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0027] Currently, the application of various smart devices (such as smart TVs, smart tablets, and smart home control screens) is becoming increasingly widespread. In some application scenarios, smart TVs and streaming media devices provide a vast amount of content, but the way users interact with the content remains very passive and simplistic.
[0028] For example, if a user wants to access extended information related to the currently viewed content (such as player data, news, etc.), they need to manually pause playback and open their phone or computer to search, a cumbersome process that interrupts the immersive experience. Alternatively, they can use the device's built-in voice assistant, but it only provides some general functions, such as searching and providing feedback on user input, and cannot deeply understand the context of specific content or provide interactive services strongly relevant to the current scenario.
[0029] It is evident that existing AI assistants can typically only provide interactive feedback based on the user's real-time input, and cannot understand the context of the corresponding content. Therefore, they cannot provide interactive services that are strongly relevant to the current scenario, affecting the interaction effect and user experience.
[0030] Specifically, in video playback scenarios, existing AI assistants cannot understand the specific plot and characters. In gaming scenarios, AI assistants cannot identify the current game state, characters, and tasks, and cannot provide genuine in-game guides or interactive features. In system settings or shopping scenarios, AI assistants cannot understand the specific products the user is currently browsing or the complex options in the settings.
[0031] Existing technologies lack a unified framework that can penetrate application barriers, understand real-time screen content at the system level, and trigger AI services, thus causing various technical problems.
[0032] Specifically, existing AI services suffer from fragmentation and scenario limitations. Current AI assistant services on smart terminals (such as smart TVs) exhibit severe functional fragmentation and scenario limitations. For example, AI functions within each application are isolated; AI in video applications cannot understand game content, and game assistants cannot recognize shopping interfaces. Furthermore, AI assistants' understanding remains superficial. Existing AI assistants (such as voice assistants on smart TVs) are mostly based on keyword matching or simple semantic understanding, unable to deeply analyze the complete visual context of the current screen. Moreover, the services provided by AI assistants are limited, offering similar services across different scenarios, lacking differentiated and in-depth services tailored to specific scenarios. For instance, when watching a basketball game, users cannot instantly access real-time statistics and player information; when playing a game, users cannot obtain targeted strategies and tips based on their current game status; and when browsing e-commerce applications, users cannot directly obtain in-depth product analysis and price comparison services.
[0033] Furthermore, there is a conflict between the transience of user interests and the non-retention of services. User interests generated while consuming digital content are instantaneous and fleeting, and current technologies lack effective mechanisms for capturing and retaining these interests. Existing technologies typically only enable one-time interactions; the results of user interactions with AI cannot be saved, requiring a complete query process to be initiated each time they are needed. User interests generated in specific scenarios (such as following a particular player or being interested in a product) cannot be converted into long-term accessible digital assets. This leads to repetitive work or calculations; for example, multiple expressions of interest in the same topic by a user require repeating the same operational steps, resulting in a waste of computing resources.
[0034] Furthermore, there are issues with broken interaction flows and inconsistent user experiences. Existing technologies suffer from severe experience disruptions in the process of accessing extended services. For example, when a user searches for content of interest, they may need to exit the current application (or pause its use) and open a browser or other application to search, interrupting the immersive experience. Switching applications may result in the loss of previously viewed information and its corresponding context. Moreover, the operation is complex, requiring multiple clicks, inputs, or voice commands to complete even simple content extension requests.
[0035] Meanwhile, there are also issues with the static nature and lack of customizability of personalized services. Existing AI services generally suffer from low personalization and rigid configurations. All users receive the exact same service in the same scenario, and the configuration is fixed, preventing users from adjusting the content and form of AI services according to their preferences. Furthermore, the service content cannot evolve; once set, it remains unchanged and cannot be optimized as user habits accumulate.
[0036] Therefore, there is a need for an integrated interactive solution on smart devices that can break down application barriers, deeply understand screen content across all scenarios, comprehend contextual information, and transform fleeting interests into customizable and evolving personalized digital assets. This solution aims to address issues such as service fragmentation, experience disruption, loss of interest, and insufficient personalization in existing technologies. Ultimately, this will transform AI services from general-purpose tools into personalized digital life assistants, upgrading them from feature stacking to scenario-based intelligence, thereby achieving a smart, seamless, and personalized user experience.
[0037] To address at least one of the aforementioned technical problems, this application proposes a smart device information processing method. Specifically, in response to an interaction trigger signal, scene data corresponding to the smart device is acquired. This scene data includes at least one of image data, audio data, text data, and metadata, whereby the metadata indicates application information corresponding to an application running on the smart device. Based on the scene data, a scene type and interaction theme information matching the scene data are determined. Based on the scene type and interaction theme information, an AI interaction service corresponding to the smart device is generated.
[0038] Thus, during the use of smart devices, in response to interaction trigger signals, scene data that is context-dependent on the current interaction content and can characterize the current scene is acquired. This determines the scene type and interaction theme information that match the scene data, and then generates AI interaction services in real time and dynamically based on the scene type and interaction theme information. In this way, the generated AI interaction service is generated based on the context information of the user-input interaction content and is strongly correlated with the current scene of the smart device. Therefore, the solution of this application can provide interactive services that are strongly correlated with the current scene, which is beneficial to improving the interaction effect and user interaction experience.
[0039] like Figure 1 As shown in the figure, this application provides a method for processing information in a smart device. Specifically, the method includes the following steps: Step S100: In response to the interaction trigger signal, scene data corresponding to the smart device is obtained, wherein the scene data includes at least one of image data, audio data, text data and metadata, and the metadata is used to indicate the application information corresponding to the application running in the smart device. Step S200: Based on the above scene data, determine the scene type and interaction theme information that match the above scene data; Step S300: Based on the above scenario type and the above interaction theme information, generate the AI interaction service corresponding to the above smart device.
[0040] It should be noted that the above-mentioned intelligent device information processing method can be executed by the intelligent device itself or by a separately set control device for controlling the intelligent device. In this embodiment of the application, the execution of the above method by the intelligent device itself is used as an example for specific explanation, but it is not intended to be a specific limitation.
[0041] Specifically, the method provided in this application embodiment can realize personalized AI interaction services based on full-scene content understanding, and the above method runs at the operating system level of smart devices.
[0042] It should be further noted that the aforementioned smart devices can be smart TVs, smart tablets, smart home control screens, etc. In this embodiment of the application, a smart TV is used as an example for specific explanation, but this is not intended as a specific limitation.
[0043] Specifically, the above method also includes: generating an interactive trigger signal in response to detecting a global trigger operation.
[0044] The aforementioned global trigger operation can be executed by the target object (i.e., the user) in any user interface as needed. For example, in some specific application scenarios, an interactive trigger signal is generated in response to the user's global trigger operation on a preset specific AI function key on a terminal device (such as a TV remote control) (this operation is effective in any interface).
[0045] Smart devices respond to interactive trigger signals by acquiring current data, such as capturing screen content (which can be done by a system-level capture module that intercepts the current screen's frame buffer data).
[0046] Specifically, the above-mentioned response to the interaction trigger signal to obtain the scene data corresponding to the smart device includes: In response to the interaction trigger signal, the screen of the smart device is captured to obtain image data, audio data is obtained according to the audio stream being played on the smart device, text data corresponding to the display interface of the smart device is identified and obtained, and metadata is determined according to the foreground application running on the smart device. The aforementioned metadata includes the package name, activity name, and window component information of the aforementioned foreground application.
[0047] Specifically, the collected image data includes full or partial screenshots displayed on the screen of the current smart device; audio data includes the audio stream being played by the system; text data includes text information on the interface obtained through the system's accessibility services or optical character recognition (OCR) technology; and the aforementioned metadata includes the package name, activity name, and window component information of the current foreground application.
[0048] Furthermore, based on the aforementioned scenario data, the scenario type and interaction theme information matching the aforementioned scenario data are determined, including: Send the above scenario data to the preset analysis platform; The analysis platform described above is used to determine the type of scenario that matches the scenario data described above. Based on the above scenario types, a target topic analysis model is determined from multiple preset topic analysis models, and interactive topic information matching the above scenario data is determined based on the target topic analysis model.
[0049] The aforementioned scenario type is one of several preset types; the aforementioned preset types include media playback scenario, game scenario, application interface scenario, and system desktop scenario; the aforementioned interaction theme information is used to characterize the interaction theme for the aforementioned scenario data under the aforementioned scenario type.
[0050] The aforementioned analysis platform can be pre-configured and adjusted according to actual needs, and no specific limitations are imposed here. In some application scenarios, an analysis engine is pre-configured as the analysis platform to achieve multi-scenario type analysis and topic extraction.
[0051] In this embodiment, the captured scene data is sent to an analysis engine for scene type identification and theme extraction. The analysis engine first determines the type of the current scene. In some application scenarios, if the current screen displays a video application playback page, the scene type is determined to be a media playback scene; if game A is currently running, it is determined to be a game scene; if the current screen displays a product details page in an e-commerce shopping software, it is determined to be an application interface scene; if the current screen displays the desktop of a smart device, it is determined to be a system desktop scene.
[0052] Furthermore, scenario-specific analysis is conducted. Based on the classification results corresponding to the above scenario types, different analysis models or strategies are invoked to perform deep semantic understanding of the scenario data, thereby determining the interactive theme information under the corresponding scenario.
[0053] In some application scenarios, if the current scenario is a media playback scenario, the media analysis model is invoked to determine the category of the currently playing media content, such as a movie or a basketball game. If it's a movie, the movie title and actor information are identified as interactive theme information; if it's a basketball game, player information is identified as interactive theme information. If the current scenario is a game scenario, the game title, current character, in-game status (e.g., being in a level boss battle), and mission objective are identified as interactive theme information. If the current scenario is an application interface scenario, key information from the current application is identified as the theme; for example, for a shopping application, product name, product category, and price are identified as interactive theme information. If the current scenario is a system desktop scenario, the application icon or user interface component (e.g., a widget) that the user is focusing on is identified as interactive theme information.
[0054] Specifically, in this embodiment, the generation and provision of AI services that are dynamic and combined with contextual information can also be realized. Specifically, based on the aforementioned scenario type and interaction topic information, one or more highly context-related AI interactive services are dynamically generated by matching from a preset service knowledge base.
[0055] It should be noted that in this embodiment, general-type services can be generated, such as "describe the current scene with AI" and "favorite this topic." Scene-specific services can also be generated for different scene types. For example, for game scenes, services such as "generate game guides," "view character skill analysis," and "start game tips discussion" can be provided; for application interface scenes, when the running application is a shopping application, services such as "compare prices," "search for similar products," and "generate product advantages and disadvantages analysis" can be provided; for system desktop scenes, services such as "quickly open applications" and "clear application cache" can be provided. It should be further noted that the specific scene type, interaction topic information, corresponding services, and preset services stored in the preset service knowledge base can all be set and adjusted according to actual needs, thereby achieving service matching and generation based on scene type and interaction topic information. The above examples are not intended to limit the scope of the service.
[0056] Furthermore, the above method also includes: in response to a save command input by the target object, generating an AI interaction theme control based on a preset format for the AI interaction service, and storing the AI interaction theme control in the data space corresponding to the target object.
[0057] In some application scenarios, the above method further includes: in response to detecting a scenario type that matches the above AI interaction service, displaying the above AI interaction theme control through the above smart device; and in response to the interaction signal triggered by the above target object based on the above AI interaction theme control, executing the above AI interaction service.
[0058] In this embodiment of the application, the AI interactive theme control is described in the form of a card, but this is not intended to be a specific limitation.
[0059] Specifically, in response to a save command input by the target (i.e., the user), the system generates a theme card in a unified format and saves it to the user's personal data space. This enables the generation and saving of personalized interactive services for each user. Furthermore, it can save cards from different scenarios, enabling the generation and saving of cross-scenario services. The card library, which stores the cards, can serve as a unified entry point beyond any single application, potentially containing cards such as "Actor A," "Strategy Guide for Boss C in Game B," and "Product D."
[0060] Based on the saved cards, scene memory and quick navigation are possible. Each card not only contains thematic information but also encapsulates the "intent" that triggers the scene. For example, clicking the "Strategy for Boss C in Game B" card will directly jump to and open Game B, and attempt to navigate to the vicinity of Boss C. Based on this solution, the ecosystem can be fully customized by users. Users can freely edit, categorize, and sort all cards, and customize AI services, such as setting up an automatic daily data news push service for the "Player E" card, thereby achieving a personalized AI service experience.
[0061] This application provides a method for processing information on a smart device. Specifically, in response to an interaction trigger signal, scene data corresponding to the smart device is acquired. The scene data includes at least one of image data, audio data, text data, and metadata. The metadata is used to indicate application information corresponding to the application running on the smart device. Based on the scene data, a scene type and interaction theme information matching the scene data are determined. Based on the scene type and interaction theme information, an AI interaction service corresponding to the smart device is generated.
[0062] Thus, during the use of smart devices, in response to interaction trigger signals, scene data that is context-dependent on the current interaction content and can characterize the current scene is acquired. This determines the scene type and interaction theme information that match the scene data, and then generates AI interaction services in real time and dynamically based on the scene type and interaction theme information. In this way, the generated AI interaction service is generated based on the context information of the user-input interaction content and is strongly correlated with the current scene of the smart device. Therefore, the solution of this application can provide interactive services that are strongly correlated with the current scene, which is beneficial to improving the interaction effect and user interaction experience.
[0063] In this application embodiment, the above-mentioned intelligent device information processing method is further described in detail based on some specific application scenarios. In some application scenarios, an AI interactive service system can be built based on the intelligent device information processing method provided in this application embodiment. Figure 2 This is a schematic diagram of an AI interactive service system architecture provided in an embodiment of this application, such as... Figure 2 As shown, the aforementioned AI interactive service system includes a user remote control, a terminal device (i.e., a smart TV), a network communication module, and a cloud-based AI service platform. Specifically, users can perform physical operations using a pre-set dedicated AI physical button on the remote control, triggering the smart TV to execute the aforementioned smart device information processing methods. The smart TV is equipped with a global triggering and capture module, a local processing and rendering module, and a theme card management module. The global triggering and capture module is used for AI button monitoring, screen frame capture, audio stream capture, and system metadata reading; the local processing and rendering module can perform data preprocessing, has a UI rendering engine, and also supports voice interaction and service command execution; the theme card management module is used for local card caching, card status synchronization, and can perform fast jumps. The smart TV communicates with the cloud-based AI service platform via the network communication module to upload scene data, receive AI service data from the cloud-based AI service platform, and synchronize card data with the cloud-based AI service platform. Network communication can be based on Hypertext Transfer Protocol Secure (HTTPS) or WebSocket, but this is not a specific limitation. The cloud-based AI service platform is equipped with a multimodal analysis engine, an AI service generation engine, user profiles, and a knowledge base. The multimodal analysis engine integrates a scene classifier, a visual recognition model, a speech recognition model, an OCR engine, and a metadata parser. The AI service generation engine includes a service rule engine, can generate natural language, and provides a content aggregation interface. The user profiles and knowledge base store user AI card libraries, user personal preference settings, and corresponding user interaction history.
[0064] Figure 3 This is a schematic flowchart illustrating a smart device information processing method provided in an embodiment of this application. Figure 3As shown, after detecting a global trigger, content is captured, followed by multimodal content analysis and theme extraction. Then, scene type identification is performed, and specificity analysis is conducted based on the identified scene type. Specifically, if the scene type is a media playback scene, media-specific analysis is performed; if it's a game scene, game-specific analysis is performed; if it's an application interface scene, application-specific analysis is performed; if it's a system desktop scene, desktop-specific analysis is performed; and for other extended scenes (i.e., scenes not specifically categorized), a general scene analysis framework can be invoked for analysis. Further, based on the analysis results, information fusion and theme entity extraction are performed, leading to dynamic AI service generation and provision. When the user performs a save operation, theme cards are created and managed, then the system returns to a standby state, awaiting the next global trigger.
[0065] Figure 4 This is a schematic diagram illustrating a specific process for multimodal content analysis and topic extraction provided in an embodiment of this application, such as... Figure 4 As shown, after receiving scene data packets, scene type determination is performed, and multimodal information recognition is conducted based on the determined scene. Specifically, in media playback scenarios, video frame analysis is performed to identify faces, objects, or scenes, and logos and text are recognized; simultaneously, audio stream analysis is performed, specifically Automatic Speech Recognition (ASR) and audio event detection. In game scenarios, game interface analysis is performed to identify the game interface and analyze character states; simultaneously, game text extraction is performed to identify game task objectives and parse state information. In application interface scenarios, OCR recognition and text semantic understanding are performed; interface layout analysis is performed to identify component types. In system desktop scenarios, icon recognition is performed to further classify applications; widget analysis is performed to identify system states. In other extended scenarios, general visual analysis can be performed to identify basic UI components; text extraction can also be performed to generate keyword themes. Additionally, metadata parsing can be performed to obtain application package names, activity names, signal source, or channel information.
[0066] Furthermore, the information analyzed above undergoes multimodal information fusion, followed by confidence level calculation and conflict resolution. For example, the confidence levels of results obtained through different identification methods are calculated separately. When different identification methods yield different results, the result with higher confidence level is used to resolve conflicts. After processing, a list of topic entities is generated, a scene context is constructed, and structured analysis results are output.
[0067] In a game scenario, a user is playing game A on a TV. The character encounters a difficult puzzle in map B and fails to solve it after multiple attempts. The user presses a dedicated AI physical button on the remote control, and the system-level capture module immediately captures the current game screen, game audio, and metadata obtained from the game progress. The scene classification engine identifies the current scene as game A and performs game-specific analysis: visual recognition identifies the game as game A, currently in map B area; interface text OCR identifies the interface prompt "Grass Element Stele Decryption"; game state analysis, based on game metadata, detects the player character repeatedly moving in the current area, determining it to be a "stuck" state; theme extracted: Game A Grass Element Stele Decryption Guide. AI service provided: The system pops up an AI service panel, providing a real-time AI description: You are currently in the Sumeru Rainforest area, attempting to solve the Grass Element Tablet puzzle.
[0068] Interactive discussion topic: Do you need me to explain the key points of solving this puzzle? In-depth services button: View text and image guides, watch video tutorials, and mark map locations with one click.
[0069] User interaction: When a user clicks to watch a video tutorial, the system will directly play the solution video for the puzzle in full screen, and automatically return to the game after the video ends.
[0070] In one media scenario, a user switches to CCTV on their set-top box to watch a live basketball game. When player C completes a spectacular dunk, the user presses the AI button. The system captures the current video frame, commentary audio, and subtitle information, performing scene analysis and theme extraction: determining it to be a media playback scenario, and further conducting multimodal analysis. Visual recognition: identifying player C as player C and team D as team D. Audio recognition: the commentary mentions "Player C scores!" Metadata analysis: combining the game's schedule information, confirming the signal source is a live sports broadcast. Themes extracted: "Player C" and "Live sports broadcast".
[0071] AI service provider: AI live description: Player C has just completed a breakthrough dunk, which is his 18th point of the game.
[0072] Interactive discussion topic: Want to know player C's dunking statistics this season? In-depth service button: View the game data, player career honors, and related highlights.
[0073] The user gives the voice command: "Save this player"; the system immediately creates a theme card for player C in "My AI Cards".
[0074] In one application interface scenario, a user is using an e-commerce application and browsing product E. While browsing product E's details page, the user presses the AI button. The system uses OCR to recognize the interface text and obtain the product name, price, and specifications. Scenario analysis and theme extraction are performed, classifying the scenario as an application interface scenario (e-commerce). Based on this scenario, specificity analysis is conducted: Text recognition: The product was identified as product E; Price extraction: The current price of 3899 yuan was identified; Context understanding: The user spent a long time on this page and scrolled through it multiple times.
[0075] Subject extracted: Product E; The system provides e-commerce specific services: AI description: This is the latest generation of gaming consoles.
[0076] In-depth services: Compare prices across the entire network, view user reviews, and compare similar products.
[0077] When a user clicks "Compare Prices Across the Web," the system displays price comparisons from other platforms.
[0078] The user then clicks the "Add" button to create a theme card for product E in "My AI Cards".
[0079] This application embodiment further illustrates the use case of a cross-scene theme card. Specifically, if a user wants to view previously saved content, they can directly enter the "AI Card" space. Unified access is possible; that is, the user can access "My AI Card" from the main TV interface, or press the AI button again to view "My AI Card" below the provided AI themes.
[0080] Cross-scenario management is achieved: the card library contains all previously created cards: "Player C" (from the media scenario); "Game A Grass Element Stele Decryption" (from the game scenario); "Product E Shopping Research" (from the e-commerce scenario).
[0081] Clicking the "Player C" card displays the player's latest match statistics and news. Clicking the "Game A Grass Element Monument Decryption" card prompts the system to ask, "Do you want to immediately jump to that location in the game?" Clicking the "Item E Shopping Research" card displays the latest price information and updated item reviews.
[0082] Furthermore, personalized customization is also possible: users can long-press the "Player C" card, select "Set Auto Update", and configure it to "Push Score Data Daily" to automatically push the score data corresponding to Player C every day.
[0083] Thus, based on this application's solution, it is possible to analyze interface content in any scenario within the smart TV system layer in real time, intelligently understand user context, and provide dynamic, scalable, and personalized AI interactive services. Specifically, it can achieve the following functions: Full-scene coverage: From games and videos to shopping, the same set of interactive logic applies to all scenarios. Deep scene understanding: Different analysis strategies are used for different scenarios to provide the most appropriate services. Interest scalability and personalization: Transforming fleeting interests into long-term manageable digital assets. Seamless experience: Through system-level global triggering, it avoids experience interruptions caused by application switching. User-driven personalized services: Users can manage and customize AI services according to their own preferences.
[0084] Figure 5 This is a schematic diagram of an AI-themed interactive details interface provided in an embodiment of this application. Specifically, the corresponding user interface of the system can provide an entry point for the AI-themed interactive details interface. After the user clicks to enter, the interactive details interface for the corresponding AI theme is displayed. Figure 5 As shown, the interface can include topics (such as basketball games), AI scene analysis descriptions, AI follow-up questions (such as game highlights), AI function services (such as game data), and AI content services (such as viewing highlights or replays). It also provides a card addition button, which allows users to add new AI cards based on the current topic.
[0085] Figure 6 This is a schematic diagram of an AI card library interface provided in an embodiment of this application. Specifically, the corresponding user interface of the system can provide an entry point for the AI card library interface. After the user clicks to enter, the AI card library interface is displayed. Figure 6 As shown, the AI card library interface displays all the AI cards stored by the user in the AI card library. The user can trigger the corresponding AI interactive service by clicking on any card.
[0086] In this embodiment, at the operating system level, in response to a global trigger operation, scene data of the current screen is captured. The scene data is then subjected to type discrimination and semantic analysis to extract the type and theme entities of the current scene. Based on the scene type and theme entities, at least one context-related AI interactive service is dynamically generated and provided. Scene types include media playback scenes, game scenes, application interface scenes, system desktop scenes, etc.; different analysis strategies are used to extract theme entities for different types of scenes. The AI interactive service is dynamically matched from the service knowledge base according to the scene type, with different service sets matched for different scene types. Furthermore, a unified format theme card is generated, which encapsulates theme entity information and intent information for quickly jumping to the scene, forming a personalized service entry set across applications and scenes. The system executing the above scheme may include: a system-level global trigger and capture module, a multimodal content analysis engine with scene classification capabilities, a dynamic AI service generation module linked to the service knowledge base, and a cross-scene theme card management module.
[0087] Based on the above solution, application barriers can be broken down, different scenarios can be intelligently distinguished, and the most appropriate AI services can be provided, upgrading it from a "TV-watching assistant" to a "digital life assistant." It can also build personal digital assets: a unified theme card library becomes a user's personal interest map and quick service center on the TV system, greatly improving device stickiness and user experience. At the same time, it provides standard interfaces for third-party developers, enabling their applications to be recognized and connected to rich AI services, forming an open AI service ecosystem.
[0088] like Figure 7 As shown, corresponding to the above-described intelligent device information processing method, this application embodiment also provides an intelligent device information processing system, which includes: The data acquisition module 710 is used to acquire scene data corresponding to the smart device in response to an interaction trigger signal. The scene data includes at least one of image data, audio data, text data and metadata. The metadata is used to indicate application information corresponding to the application running in the smart device. The data processing module 720 is used to determine the scene type and interaction theme information that match the scene data based on the scene data mentioned above. The service generation module 730 is used to generate AI interaction services corresponding to the smart devices based on the above-mentioned scenario type and interaction theme information.
[0089] Thus, during the use of smart devices, in response to interaction trigger signals, scene data that is context-dependent on the current interaction content and can characterize the current scene is acquired. This determines the scene type and interaction theme information that match the scene data, and then generates AI interaction services in real time and dynamically based on the scene type and interaction theme information. In this way, the generated AI interaction service is generated based on the context information of the user-input interaction content and is strongly correlated with the current scene of the smart device. Therefore, the solution of this application can provide interactive services that are strongly correlated with the current scene, which is beneficial to improving the interaction effect and user interaction experience.
[0090] It should be noted that the specific structure and implementation of the above-mentioned intelligent device information processing system and its various modules or units can be referred to the corresponding descriptions in the above method embodiments, and will not be repeated here.
[0091] It should be noted that the division of the various modules of the above-mentioned intelligent device information processing system is not unique and is not intended as a specific limitation.
[0092] Based on the above embodiments, this application also provides a terminal, the principle block diagram of which can be as follows: Figure 8 As shown. The terminal includes a processor, memory, network interface, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps of any of the above-described intelligent device information processing methods. The display screen can be a liquid crystal display (LCD) or an e-ink display.
[0093] Those skilled in the art will understand that Figure 8 The block diagram shown is only a partial structural diagram related to the solution of this application and does not constitute a limitation on the terminal on which the solution of this application is applied. The specific terminal may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0094] In one embodiment, a terminal is provided, the terminal including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of any of the smart device information processing methods provided in the embodiments of this application.
[0095] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the intelligent device information processing methods provided in this application.
[0096] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0098] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0099] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0100] In the embodiments provided in this application, it should be understood that the disclosed systems / terminal devices and methods can be implemented in other ways. For example, the system / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units described above is merely a logical functional division, and in actual implementation, it can be divided in other ways. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0101] If the integrated modules / units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, and software distribution media, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0102] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions are not in essence a departure from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for processing information in intelligent devices, characterized in that, The method includes: In response to an interaction trigger signal, scene data corresponding to the smart device is acquired, wherein the scene data includes at least one of image data, audio data, text data, and metadata, and the metadata is used to indicate application information corresponding to the application running in the smart device; Based on the scene data, determine the scene type and interaction theme information that match the scene data; Based on the scenario type and the interaction theme information, generate the AI interaction service corresponding to the smart device.
2. The intelligent device information processing method according to claim 1, characterized in that, The method further includes: In response to the detection of a global trigger operation, an interactive trigger signal is generated.
3. The intelligent device information processing method according to claim 1, characterized in that, The step of acquiring scene data corresponding to the smart device in response to the interaction trigger signal includes: In response to an interaction trigger signal, the system takes a screenshot of the smart device to obtain image data, obtains audio data based on the audio stream being played on the smart device, identifies and obtains text data corresponding to the display interface of the smart device, and determines metadata based on the foreground application running on the smart device. The metadata includes the package name, activity name, and window component information of the foreground application.
4. The intelligent device information processing method according to claim 1, characterized in that, The step of determining the scene type and interaction theme information matching the scene data based on the scene data includes: The scene data is sent to a preset analysis platform; The analysis platform is used to determine the scene type that matches the scene data. Based on the scenario type, a target topic analysis model is determined from a set of preset topic analysis models, and interactive topic information matching the scenario data is determined based on the target topic analysis model.
5. The intelligent device information processing method according to claim 4, characterized in that, The scene type is one of several preset types; The various preset types include media playback scenarios, game scenarios, application interface scenarios, and system desktop scenarios; The interaction theme information is used to characterize the interaction theme for the scene data under the scene type.
6. The intelligent device information processing method according to any one of claims 1 to 5, characterized in that, The method further includes: In response to the save command input by the target object, an AI interaction theme control based on a preset format is generated for the AI interaction service, and the AI interaction theme control is stored in the data space corresponding to the target object.
7. The intelligent device information processing method according to claim 6, characterized in that, The method further includes: In response to detecting a scene type that matches the AI interaction service, the AI interaction theme control is displayed on the smart device; In response to the interaction signal triggered by the target object based on the AI interactive theme control, the AI interaction service is executed.
8. An intelligent device information processing system, characterized in that, The system includes: The data acquisition module is used to acquire scene data corresponding to the smart device in response to an interaction trigger signal. The scene data includes at least one of image data, audio data, text data, and metadata. The metadata is used to indicate application information corresponding to the application running in the smart device. The data processing module is used to determine the scene type and interaction theme information that match the scene data based on the scene data. The service generation module is used to generate AI interaction services corresponding to the smart device based on the scenario type and the interaction theme information.
9. A terminal, characterized in that, The terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the intelligent device information processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the intelligent device information processing method as described in any one of claims 1 to 7.