Scene setting method, apparatus and device of intelligent device and storage medium

By acquiring multi-dimensional data from smart devices and performing feature analysis, main scene tags and sub-scene tags are generated, solving the problem of the inability to dynamically adjust the environmental atmosphere in existing technologies and realizing a more immersive smart home linkage control experience.

CN122632682APending Publication Date: 2026-08-25SHENZHEN TCL NEW-TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610704634.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing smart home linkage control methods cannot dynamically adjust the environmental atmosphere according to the deep characteristics of the content, making it difficult to meet users' potential demand for a deep immersive experience that is highly integrated with the content context.

Method used

By acquiring multi-dimensional data, including image data, audio data, and metadata from the main control device, as well as object data from the viewed object, feature extraction and feature analysis are performed to generate main scene tags and sub-scene tags. Based on these tags, the scene setting content is determined, and the smart device is controlled to set the scene.

Benefits of technology

It enables the dynamic adjustment of multiple smart devices based on the content and plot, creating an immersive atmosphere that is highly synchronized with the emotions of the visuals, thus improving the intelligence of scene setting and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122632682A_ABST
    Figure CN122632682A_ABST
Patent Text Reader

Abstract

The application discloses a scene setting method and device of an intelligent device, equipment and a storage medium, and is used for improving the intelligent dimension of scene setting and making the scene setting have a more immersive experience. The method comprises the following steps: acquiring multi-dimensional data, wherein the multi-dimensional data comprises image data, audio data and metadata of playing content currently played by a master device, object data of a current viewing object, and the metadata is used for describing description information of the playing content; performing feature extraction and feature analysis on the multi-dimensional data to obtain a main scene label and a sub-scene label, wherein the sub-scene label is a sub-label of the main scene label; determining scene setting content based on the main scene label and the sub-scene label, wherein the scene setting content comprises one or more control instructions; and controlling one or more intelligent devices based on the scene setting content to complete scene setting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent devices, and particularly to a method, device, equipment and storage medium for scene setting of intelligent devices. Background Art

[0002] With the rapid development of Internet of Things and artificial intelligence technologies, smart home systems have evolved from the remote control of single devices to the scene-based intelligent stage of multi-device collaborative linkage. Through preset or user-defined scene modes, such as "leaving home mode", "sleep mode", etc., the system can integrate and control multiple home devices such as lights, curtains, air conditioners, etc., greatly improving the living convenience and comfort of users. Among them, the home entertainment scene, especially the viewing experience around large-screen display devices such as TVs and projectors, has become an important application field of smart home linkage control.

[0003] Currently, the linkage for the viewing scene usually adopts a fixed strategy based on event triggering. For example, the system detects the power-on signal of the TV or the user manually activates the "viewing mode" to trigger a set of pre-configured instructions, such as turning off or dimming the lights in the area where the TV is located, closing the curtains, etc., so as to create a suitable basic environment for users to watch movies. This method simplifies the user's operation to a certain extent and provides a preliminary automated experience compared with completely manual adjustment.

[0004] However, the current triggering logic of linkage control only depends on the macroscopic state of the device (such as on / off), and is completely decoupled from the specific content played by the device. That is, the current linkage control method cannot perceive and respond to the deep features of the content, such as the ups and downs of the plot, the color tone of the picture, the atmosphere of the scene, and the dynamics of the sound effect. Therefore, it is difficult to dynamically and finely adjust the environmental atmosphere according to the content, and it is difficult to meet the potential needs of users for a deep immersive experience highly integrated with the content scenario. The intelligent dimension of scene linkage and the depth of user experience both need to be further improved. Summary of the Invention

[0005] The embodiments of this application provide a method, device, equipment and storage medium for scene setting of intelligent devices, which are used to improve the intelligent dimension of scene setting and make the scene setting more immersive.

[0006] The technical solutions adopted by the present invention to solve the problems are as follows: In a first aspect, this application provides a method for scene setting of intelligent devices, including: Obtain multi-dimensional data, where the multi-dimensional data includes image data, audio data and metadata of the playback content currently played by the master device, and object data of the current viewing object, and the metadata is used to describe the description information of the playback content; Feature extraction and feature analysis are performed on the multi-dimensional data to obtain the main scene label and sub-scene label, where the sub-scene label is a sub-label of the main scene label; The scene settings content is determined based on the main scene label and the sub-scene label, and the scene settings content includes one or more control commands; Based on the scenario settings, one or more smart devices can be controlled to complete the scenario setup.

[0007] In some embodiments of this application, obtaining multi-dimensional data includes: The image data, audio data, and metadata are obtained based on the target interface between the main control device and the content platform. The main control device uses its camera and / or microphone to acquire object data of the currently viewed object.

[0008] In some embodiments of this application, feature extraction and feature analysis are performed on the multi-dimensional data to obtain main scene labels and sub-scene labels, including: Based on a multi-feature weighted algorithm, feature extraction and feature analysis are performed on the multi-dimensional data to obtain the main scene label and the sub-scene label; And / or, Based on the intent recognition model, feature extraction and feature analysis are performed on the multi-dimensional data to obtain the main scene label and the sub-scene label.

[0009] In some implementations of this application, determining the scene settings based on the main scene label and the sub-scene label includes: The main scene tag and the sub-scene tag are matched in the scene recommendation template library to obtain the scene setting content; and / or; The main scene label and the sub-scene label are processed based on the processing model to obtain the scene setting content.

[0010] In some embodiments of this application, after determining the scene setting content based on the main scene label and the sub-scene label, the method further includes: The main control device displays a first interface on its screen. This first interface is used to display the scene settings and a first interactive control. The first interactive control includes a confirmation control, a parameter adjustment control, and a rejection control. The confirmation control indicates that the scene settings are being applied. The parameter adjustment control indicates that the scene settings are being updated. The rejection control indicates that the application of the scene settings is being interrupted.

[0011] In some embodiments of this application, after controlling one or more smart devices based on the scene settings to complete the scene settings, the method further includes: The main control device displays a second interface on its screen. This second interface is used to display a second interactive control, which is used to collect feedback data from the currently viewing object regarding the scene settings.

[0012] In some embodiments of this application, the method further includes: Obtain feedback information, which includes the interaction data and feedback data of the currently viewed object. The interaction data is used to characterize the interruption operation, application operation and / or parameter adjustment operation of the currently viewed object on the scene setting content. The feedback data is used to characterize the feedback of the currently viewed object on the scene setting after the scene setting content is made. Based on this feedback information, the scenario recommendation template library or processing model is updated, and this scenario recommendation template library or processing model is used to generate the scenario setting content.

[0013] Secondly, this application provides a scene setting device for a smart device, comprising: The acquisition module is used to acquire multi-dimensional data, which includes image data, audio data and metadata of the playback content currently being played by the main control device, and object data of the currently viewed object. The metadata is used to describe the descriptive information of the playback content. The processing module is used to extract and analyze features from the multi-dimensional data to obtain a main scene label and a sub-scene label, wherein the sub-scene label is a sub-label of the main scene label; based on the main scene label and the sub-scene label, the scene setting content is determined, which includes one or more control commands; based on the scene setting content, one or more smart devices are controlled to complete the scene setting.

[0014] Thirdly, this application also provides a computer device, which includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the methods of any one of the first aspects.

[0015] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the method of any one of the first aspects.

[0016] The beneficial effects of this invention are as follows: By combining image data, audio data, and metadata of the playback content, as well as object data of the viewing object, scene recognition is performed to obtain main scene tags and sub-scene tags. Scene setting content is then generated based on these tags. This scene recognition incorporates multi-dimensional reference data, enhancing the intelligence of scene setting. Furthermore, the combined use of main and sub-scene tags allows for smart device interaction beyond a single main scene mode, introducing more specific content and plot tags. This enables the generated scene setting content to dynamically adjust multiple smart devices according to the content and plot, creating an immersive atmosphere highly synchronized with the emotional content of the visuals, resulting in a more immersive and engaging scene setting experience. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of an embodiment of the scene setting method for smart devices provided in this invention; Figure 3a This is a schematic diagram of an embodiment of the interactive interface provided in this invention; Figure 3b This is a schematic diagram of another embodiment of the interactive interface provided in this invention; Figure 3c This is a schematic diagram of another embodiment of the interactive interface provided in this invention; Figure 4 This is a flowchart illustrating a scene setting method for smart devices provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of an embodiment of the scene setting device for intelligent devices provided in this invention. Figure 6 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features.

[0021] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0022] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.

[0023] With the rapid development of Internet of Things and artificial intelligence technologies, smart home systems have evolved from the remote control of single devices to the scenario-based intelligence stage of multi-device collaborative linkage. Through preset or user-defined scenario modes, such as "away mode", "sleep mode", etc., the system can integrate and control multiple home devices, such as lights, curtains, air conditioners, etc., greatly enhancing the convenience and comfort of users' living. Among them, the home entertainment scenario, especially the viewing experience around large-screen display devices such as TVs and projectors, has become an important application area for smart home linkage control. Currently, the linkage for the viewing scenario usually adopts a fixed strategy based on event triggering. For example, the system detects the power-on signal of the TV or the user manually activates the "viewing mode" to trigger a set of pre-configured instructions, such as turning off or dimming the lights in the area where the TV is located, closing the curtains, etc., so as to create a suitable basic environment for users to watch movies. This method simplifies the user's operation to a certain extent and provides a preliminary automated experience compared with completely manual adjustment.

[0024] However, the current trigger logic of linkage control only depends on the macroscopic state of the device (such as on / off), and is completely decoupled from the specific content played by the device. That is, the current linkage control method cannot perceive and respond to the deep features of the content, such as the ups and downs of the plot, the tone of the picture, the atmosphere of the scene, and the dynamics of the sound effect. Therefore, it is difficult to dynamically and finely adjust the environmental atmosphere according to the content, and it is difficult to meet the potential needs of users for a deep immersive experience that is highly integrated with the content scenario. The intelligent dimension of the scene linkage and the depth of the user experience both need to be further improved.

[0025] To address this technical problem, this application provides the following technical solution: Acquiring multi-dimensional data, including image data, audio data, and metadata of the currently playing content on the main control device, and object data of the currently viewed object, whereby the metadata describes the content being played; performing feature extraction and analysis on the multi-dimensional data to obtain a main scene tag and sub-scene tags, where the sub-scene tags are sub-tags of the main scene tag; determining scene setting content based on the main scene tag and the sub-scene tag, whereby the scene setting content includes one or more control commands; and controlling one or more smart devices based on the scene setting content to complete the scene setting. By combining the image data, audio data, and metadata of the playing content, as well as the object data of the viewed object, scene recognition is performed to obtain the main scene tag and sub-scene tag, and scene setting content is generated based on the main scene tag and sub-scene tag. This scene recognition introduces multi-dimensional reference data, improving the intelligence of scene setting. Simultaneously, the combined use of the main scene tag and sub-scene tag allows the linkage of smart devices to break through the single main scene mode, introducing more specific content plot tags. This allows the generated scene settings to dynamically adjust multiple smart devices according to the plot, creating an immersive atmosphere that is highly synchronized with the emotions of the visuals, making the scene settings more immersive and enhancing the rendering experience.

[0026] This application provides a method, apparatus, device, and storage medium for setting scenes on a smart device, which enhances the intelligence of scene settings and makes them more immersive. The electronic device provided in this application can be implemented as various types of user terminals or as a server.

[0027] Electronic devices can improve the intelligence dimension of scene setting by running the scene setting method of the intelligent device provided in the embodiments of this application, so as to make the scene setting more immersive.

[0028] The above methods can be applied to many smart devices, such as smart TVs, smart speakers, smart air conditioners, and so on.

[0029] In one exemplary solution, the scene setting method of the smart device can be applied to the linkage control of smart TVs and other smart home devices. For example, when a smart TV is playing video content, it can obtain the video content's image data, audio data, and metadata (which may include the video content's name, release date, synopsis, and category tags) through the corresponding interface of the content platform to which the video content belongs. It can also collect object data of the currently viewed object (which includes the object's image data and / or audio data) through a camera and / or microphone connected to the smart TV. Then, the smart TV uses a built-in intent recognition model to perform scene recognition on the video content's image data, audio data, and metadata, as well as the object data of the viewed object, to obtain a main scene tag and sub-scene tags (e.g., the main scene tag is "Movie Mode - Action Movie," and the sub-scene tag is "Intense Chase / Gunfight"). Based on the main scene tag and the sub-scene tag, the smart TV generates scene setting content (e.g., the whole room lighting simulates flashing blue and red police lights, light strips light up instantaneously in conjunction with explosion scenes, the air conditioning temperature is slightly lowered, and the fan is turned on to create a flurry of air). Finally, the smart TV distributes the control commands in the scene setting content to the corresponding smart devices to complete the scene setting.

[0030] It should be understood that the above is only an exemplary application scenario of the scene setting method for smart devices. There are many other specific application scenarios, which are not limited here.

[0031] The scene setting method for smart devices provided in this application embodiment is applied to, for example, Figure 1 The system architecture diagram shown is for your reference. Figure 1 To support the scene setting method for a smart device, the terminal device 100 connects to the server 300 via network 200, and the server 300 connects to the database 400. Network 200 can be a wide area network (WAN), a local area network (LAN), or a combination of both. The client used to implement the scene setting scheme for the smart device is deployed on the terminal device 100, or it can run on the terminal device 100 as a standalone application. The specific form of the client is not limited here.

[0032] The server 300 involved in this application can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.

[0033] Terminal equipment 100, also known as user equipment (UE), mobile station (MS), mobile terminal (MT), customer premises equipment (CPE), etc., can be a device that includes both receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such equipment can include cellular or other communication devices with single-line displays, multi-line displays, or no multi-line displays. Examples include handheld devices with wireless connectivity, vehicle-mounted devices, and machine-type communication (MTC) terminals. Currently, terminal devices can include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving cars, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, and wireless terminals in smart homes. For example, wireless terminals in self-driving cars can be drones, helicopters, or airplanes. For example, wireless terminals in vehicle-to-everything (V2X) systems can be in-vehicle equipment, vehicle-mounted equipment, in-vehicle modules, vehicles, or ships. Wireless terminals in industrial control can be cameras, robots, or robotic arms. Wireless terminals in smart homes can be televisions, air conditioners, robot vacuums, speakers, or set-top boxes.

[0034] It should be noted that the terminal device may be a device or apparatus with a chip, or a device or apparatus with integrated circuitry, or a chip, module or control unit in the device or apparatus shown above. This application does not limit the specific device.

[0035] The solution provided in this application can be completed by the cooperation of terminal device 100 and server 300.

[0036] In short, a database can be viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, shared by multiple users, with minimal redundancy, and independent of application programs. A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or Extensible Markup Language (XML); or according to the type of computer they support, such as server clusters or mobile phones; or according to the query language used, such as Structured Query Language (SQL) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, supporting multiple query languages ​​simultaneously. In this application, database 400 can be used to store image data, audio data and metadata of the playback content, object data of the viewing object, intent recognition model, processing model and other data.

[0037] Those skilled in the art will understand that Figure 1 The system architecture diagram shown is one possible system architecture for this application and does not constitute a limitation on the system architecture of this application. Other system architectures may include more advanced architectures. Figure 1 The number of more or fewer terminal devices or servers shown, for example Figure 1 The diagram shows one server. It is understood that the system architecture may also include one or more other terminal devices or servers, which are not limited here.

[0038] It should be noted that, Figure 1 The system architecture shown is an example. The servers and scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of servers and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0039] like Figure 2The diagram shown is a flowchart of an embodiment of the scene setting method for a smart device in this application. The following description uses a smart TV as the execution subject to illustrate the scene setting method for the smart device, which may include the following steps 201-204, as detailed below: 201. Obtain multi-dimensional data, which includes image data, audio data, and metadata of the content currently being played on the main control device, as well as object data of the currently viewed object. The metadata is used to describe the descriptive information of the content being played.

[0040] When the smart TV is the main control device, it has applications corresponding to various content platforms installed on it. In response to the user's operation, the smart TV starts playing the corresponding video or audio (i.e., the content played by the main control device). Then, the smart TV can call the interface of the content platform to which the content belongs to obtain the image data, audio data, and metadata of the content. At the same time, it can call the smart TV's camera and / or microphone to obtain the object data of the currently viewed object. The image data, audio data, and metadata of the content, as well as the object data of the currently viewed object (the object data may include the image data and / or audio data of the viewed object, where the audio data may be a specific voice command, non-specific voice feedback, or background sound in the current environment) serve as the multi-dimensional data.

[0041] Optionally, the smart TV can perform the multi-dimensional data collection task based on the current status and / or a preset collection strategy.

[0042] When the smart TV initiates a multi-dimensional data collection task, the event trigger can be a key event in the playback operation. For example, the start of playback, the viewer issuing a voice command, the start or end of an advertising period, the arrival of a preset periodic detection time, etc.

[0043] After triggering the multi-dimensional data collection task, the smart TV can select a suitable collection scheme from the strategy library (including the preset collection strategy) based on the type of triggering event and the current context (e.g., whether a movie or news is playing, whether it is daytime or nighttime). For example, when playing a movie, high-frequency audio and video content analysis and user expression analysis can be enabled; when playing news, only metadata and news keyword information can be collected.

[0044] After determining the acquisition strategy, the smart TV can generate specific acquisition instructions. For example: "Request the audio and video features of content ID mov123 at timestamp 01:35:10" or "Start the camera to detect the number of viewers and their attention span".

[0045] Optionally, in practical applications, this smart TV can also adjust its data collection strategy based on system resource usage (such as CPU usage and network bandwidth usage) to avoid affecting the user's normal viewing experience due to data collection. Alternatively, the smart TV can dynamically adjust its data collection strategy based on privacy settings (such as when the user turns off the camera).

[0046] Based on the above description, the process by which the smart TV collects image data, audio data, and metadata of the content it plays is explained below: In one exemplary solution, the smart TV can generate a content identifier and timestamp for the content being played; then, based on the content identifier and timestamp, it requests the metadata of the content (such as title, type (such as movie, TV series, short drama, documentary, sports event, news broadcast, etc.), tags (such as science fiction, suspense, romance), chapter information, advertising time points, and opening and closing credits time points, etc.) as well as the image data (in this case, the image data can be the keyframe image corresponding to the timestamp) and audio data (the audio data can be sound effects, such as explosion sounds, wind sounds, rain sounds, ocean waves sounds, etc.; or the audio data can be voice data, such as actors' dialogue, documentary narration, etc.) from the server of the content platform to which the content belongs.

[0047] Optionally, to improve data processing efficiency, the smart TV can also directly obtain the image feature vector or audio feature vector corresponding to the timestamp from the content platform. In this case, the image feature vector is used as the image data, and the audio feature vector is used as the audio data.

[0048] The following explains the process by which the smart TV collects object data from the viewers: In one exemplary solution, the smart TV can call the camera API provided by the TV operating system (the camera can be a built-in camera of the smart TV or an external camera of the smart TV) to obtain a video stream, which serves as the image data of the object data.

[0049] The smart TV can call the audio API provided by the TV operating system (which can be the microphone built into the smart TV or the microphone on the remote control associated with the smart TV) to collect ambient sound (which includes the user's specific voice commands, non-specific voice feedback, and background noise), and use the ambient sound as the audio data of the object data.

[0050] 202. Perform feature extraction and feature analysis on the multi-dimensional data to obtain the main scene label and sub-scene label, where the sub-scene label is a sub-label of the main scene label.

[0051] After acquiring the aforementioned multi-dimensional data, the smart TV can perform feature extraction and feature analysis on the multi-dimensional data to obtain the main scene label and sub-scene label. The sub-scene label is a sub-label of the main scene label (or a secondary label; for example, if the main scene label is "Movie Mode - Action Movie", the sub-scene label could be "Intense Chase / Gunfight").

[0052] In one exemplary solution, the smart TV can perform feature extraction and feature analysis on the multi-dimensional data based on a multi-feature weighted algorithm to obtain the main scene label and the sub-scene label.

[0053] In this calculation process, the multi-feature weighting algorithm can be a static weighting algorithm, a dynamic weighting algorithm, or a machine learning weighting algorithm.

[0054] For example, when the multi-feature weighting algorithm is a dynamic weighting algorithm, the smart TV can determine the main scene label of the content being played based on its metadata; then, based on the main scene label, it can determine the corresponding multi-feature weighting parameters from a preset weighting parameter library; and then, based on the multi-feature weighting parameters, it can perform weighted calculations on the multi-dimensional data corresponding to the content being played to obtain the sub-scene label. For example, when the main scene label is "action movie," an "action movie weight table" can be loaded, which can be represented as follows: explosion sound weight 0.5, screen shake weight 0.3, fast-paced music weight 0.2. When the main scene label is "romance movie," a "romance movie weight table" is loaded, which can be represented as follows: warm color tone weight 0.4, soft music weight 0.4, character eye contact weight 0.2.

[0055] When playing "Fast & Furious," the system uses the first set of weights, which easily identifies the "car chase and gunfight" sub-scene tag. When playing "Titanic," the system switches to the second set of weights, making it easier to identify the "romantic confession" sub-scene tag.

[0056] In another exemplary solution, the smart TV can perform feature extraction and feature analysis on the multi-dimensional data based on an intent recognition model to obtain the main scene label and the sub-scene label.

[0057] In this calculation process, the smart TV can first perform feature processing on the multi-dimensional data to obtain feature vectors, and then use the trained intent recognition model to recognize the feature vectors to obtain the main scene label and sub-scene label.

[0058] Optionally, the smart TV can perform feature processing and feature analysis on this multi-dimensional data in the following ways: For the image data of the playback content, a lightweight computer vision model can be run locally to extract features from the image data, thereby obtaining image features. These image features include low-level image features and high-level image features. The low-level image features can include color histograms (to determine hue, such as cool or warm colors), brightness, and contrast. The high-level image features can include object recognition features (such as identifying objects in the image), scene classification features (identifying scenes as indoors, outdoors, cities, natural scenery, etc.), facial and facial expression features (emotional features of people in the video, such as happiness or sadness), and motion features (such as chasing).

[0059] For the audio data of the played content, a lightweight speech recognition model can be run locally to extract features from the audio data to obtain audio features. These audio features can include low-level audio features and high-level audio features. The low-level audio features can include pitch, volume, speech rate, rhythm, etc. The high-level audio features can include keyword features obtained from converting speech to text, voiceprint features, audio event detection features (such as sound effects, explosions, car engine sounds, applause, etc.), and music analysis features (such as background music type and rhythm, for example, soothing music with a slow tempo).

[0060] For the metadata of the content being played, a lightweight text analysis model can be run locally to extract text features. For example, word segmentation and sentiment analysis can be performed directly on the title and description to extract keywords.

[0061] Run a lightweight computer vision model locally to extract object features such as "number of viewers", "face position", "direction of gaze" and "approximate posture (e.g., sitting upright, lying down)".

[0062] Through local model analysis, specific user commands, non-voice feedback sounds (such as laughter, exclamations), or environmental conditions (such as quiet, noisy, or baby crying) can be identified.

[0063] After completing the above feature extraction, the image features of the playback content, the audio features of the playback content, the text features of the metadata of the playback content, the image features of the viewing object, and the audio features of the viewing object can be weighted and analyzed to obtain the main scene label and the sub-scene label.

[0064] Alternatively, the image features of the playback content, the audio features of the playback content, the text features of the metadata of the playback content, the image features of the viewing object, and the audio features of the viewing object can be fused together and then input into the intent recognition model for recognition to obtain the main scene label and the sub-scene label.

[0065] In one exemplary solution, to more effectively identify the main scene label and the sub-scene label, they can be identified separately. For example, when the metadata containing the playback content is available, since this metadata may include the classification information of the playback content, the main scene label can first be identified based on the metadata using rules or a simple classification model. Then, a sub-scene identification model is loaded based on the main scene label, and sub-scene identification is performed based on the feature vectors corresponding to the multi-dimensional data to obtain the sub-scene label.

[0066] 203. Determine the scene settings content based on the main scene label and the sub-scene label. The scene settings content includes one or more control commands.

[0067] After the smart TV obtains the main scene tag and sub-scene tag corresponding to the content being played, it can use various methods to determine the scene settings.

[0068] In one exemplary solution, the smart TV matches the main scene tag and the sub-scene tag from the scene recommendation template library to obtain the scene setting content. In this solution, the smart TV first queries the scene recommendation template library based on the main scene tag to obtain the basic template corresponding to the main scene tag (it should be understood that this template can be understood as a set of general effect descriptors); then it queries based on the sub-scene tag to obtain the enhanced template corresponding to the sub-scene tag (usually this enhanced template is used to modify or supplement the basic template); finally, the basic template and the enhanced template are merged to obtain a target template; then, based on the target template, the roles and capabilities of configurable smart devices are queried (for example, configurable smart devices can be queried through the device registry, where the device registry can be pre-registered by the user. For example, when adding a smart device to the home, the user can add the smart device to the device registry and note its role and capabilities. For example, adding a smart speaker can register the smart speaker as an audio player and allow it to connect to the smart TV for audio playback); then, based on the effect description of the target template and the roles and capabilities of each configurable smart device, control instructions corresponding to each configurable smart device are generated, and these control instructions serve as the scene setting content. For example, if the main scene is labeled "Movie Playback - Movie" and the sub-scene is labeled "Intense Gunfight," the target template could be represented as "Lighting effect: Dynamic cool-colored flashing; Audio effect: Bass enhancement and surround sound; Environmental effect: No interference." Assuming the configurable smart devices include living room light strips, a sofa floor lamp, a living room soundbar, curtains, and a smart door lock, the generated scene settings could be as follows: The living room light strip has its special effects enabled and is set to cool-colored flashing; the sofa floor lamp is not included in the settings; the living room soundbar is set to surround mode and the mode is set to "Movie"; the living room soundbar's bass level is also set to 90; and the living room curtains are set to closed.

[0069] In another exemplary solution, the smart TV processes the main scene label and the sub-scene label based on a processing model to obtain the scene setting content. In this solution, the processing model can be understood as a pre-trained scene setting control instruction generation model, whose input data are various scene labels and information on configurable smart devices, and whose output data are the predicted control instructions corresponding to the configurable smart devices.

[0070] Optionally, after generating the scenario settings, in order to empower users with the right to know and intervene in system commands, and to support personalized fine-tuning, thereby avoiding the problem of users lacking control due to the inability to know the specific operations in automated linkage, this embodiment may also provide the following solution: The smart TV displays a first interface on its screen, which is used to display the scene settings and a first interactive control. The first interactive control may include a confirmation control, a parameter adjustment control, and a rejection control. The confirmation control indicates that the scene settings are applied, the parameter adjustment control indicates that the scene settings are updated, and the rejection control indicates that the application of the scene settings is interrupted.

[0071] In one exemplary solution, the first interface can be as follows: Figure 3a As shown, the scene settings can be represented as: "Living room light strip on with special effects, set to cold light flashing; sofa floor lamp not included in the settings; living room soundbar selected for surround mode, mode set to movie; living room soundbar also set to bass level 90; living room curtains set to closed." Its interactive controls include: "Confirm," "Deny," and "Adjust." When the user clicks "Confirm," the scene settings displayed on the current first interface are applied directly; when the user clicks "Deny," the scene settings displayed on the current first interface are rejected, and the scene settings are regenerated or no further scene settings are applied; when the user clicks "Adjust," the first interface can be changed from... Figure 3a The changes shown are as follows: Figure 3b As shown. In Figure 3b In this mode, the scene settings can be displayed in an adjustable manner. This means you can select the corresponding controls to make adjustments. For example, you can change the sofa floor lamp from off to on, and set the effect to a blinking mode.

[0072] 204. Based on the scenario settings, control one or more smart devices to complete the scenario settings.

[0073] After acquiring the scene settings, the smart TV generates corresponding control commands based on the interaction protocol of one or more smart devices and distributes these commands to the corresponding smart devices to complete the scene settings. For example, after receiving the command "Turn on special effects for the living room light strip, set the effect to cold light flashing," the "XX light strip adapter" translates it into a specific command conforming to the XX protocol and sends it to the light strip via the wireless network. Similarly, after receiving the command "Select surround mode for the living room soundbar, set the mode to movie; also set the bass level to 90," the "YY speaker adapter" may translate it into a request conforming to a specific or proprietary protocol and send it to the speaker.

[0074] In this embodiment, to enhance the transparency, flexibility, and user satisfaction of human-computer interaction, an interactive interface can be provided to collect user feedback. In one exemplary solution, the smart TV can display a second interface on its screen after a preset time period has elapsed since the scene settings were applied. This second interface displays a second interactive control, which is used to collect feedback data from the currently viewing user regarding the scene settings. For example, such as... Figure 3c As shown, a pop-up window appears on the smart TV screen, displaying the question "Keep current scene settings?" and two selection controls: "Yes" and "No". If the user selects "Yes", the current scene settings can be applied; if the user selects "No", the normal settings can be restored or the scene settings can be regenerated and applied again.

[0075] Optionally, the scene settings can be generated periodically; or generated in real time based on key scenes of the playback content; or generated by user commands. No specific restrictions are imposed here.

[0076] In this embodiment, to further enhance the intelligence of scene settings and user satisfaction, feedback information can be obtained, and the scene recommendation template library or processing model can be updated and optimized based on this feedback information. The feedback information includes the interaction data and feedback data of the currently viewing object. The interaction data characterizes the currently viewing object's interruption, application, and / or parameter adjustment operations on the scene setting content, while the feedback data characterizes the currently viewing object's feedback after setting the scene based on the scene setting content.

[0077] As described above, this embodiment combines image data, audio data, and metadata of the playback content, as well as object data of the viewing object, to perform scene recognition, thereby obtaining main scene tags and sub-scene tags. Scene setting content is then generated based on these tags. This scene recognition introduces multi-dimensional reference data, enhancing the intelligence of scene setting. Furthermore, the combined use of main and sub-scene tags allows the linkage of smart devices to transcend the single main scene mode, introducing more specific content plot tags. This enables the generated scene setting content to dynamically adjust multiple smart devices according to the content plot, creating an immersive atmosphere highly synchronized with the emotional content of the visuals, resulting in a more immersive rendering experience. On the other hand, providing an interactive interface significantly enhances the transparency, flexibility, and user satisfaction of human-computer interaction. Moreover, providing model optimization schemes can further improve the intelligence of scene setting and user satisfaction.

[0078] Based on the above description, the following will be used as an example. Figure 4 The processing flow shown illustrates the scene setting method for the smart device provided in this application: As shown in Figure 4 Figure [not provided], the smart TV, as the master device, can collect multi-dimensional features, including picture features, audio features, and metadata features. Then, based on the multi-feature weighting algorithm, feature analysis is performed on the picture features, audio features, and metadata features to obtain scene tags. Then, scene setting content is generated based on the scene tags. Then, interaction confirmation is carried out with the viewing user through the interactive interface displayed on the smart TV. For example, the application of the scene setting content can be determined through the confirmation control, and then the smart TV distributes instructions, and the smart device executes the instructions to complete the scene setting. The adjustment of the scene setting content can be carried out through the adjustment control, and then the smart TV distributes the adjusted instructions, and the smart device executes the adjusted instructions to complete the scene setting. The execution of the scene setting content can be interrupted through the rejection control. The smart TV collects various user feedbacks and then optimizes the model to adjust the generation of the scene setting content.

[0079] To better implement the scene setting method of the smart device in the embodiments of the present application, based on the scene setting method of the smart device, an apparatus for scene setting of a smart device is further provided in the embodiments of the present application. As shown in Figure 5 Figure [not provided], the apparatus 500 for scene setting of the smart device includes: An acquisition module 501, configured to acquire multi-dimensional data, where the multi-dimensional data includes image data, audio data, and metadata of the playback content currently played by the master device, and object data of the current viewing object, and the metadata is used to describe the description information of the playback content; A processing module 502, configured to perform feature extraction and feature analysis on the multi-dimensional data to obtain a main scene tag and a sub-scene tag, where the sub-scene tag is a sub-tag of the main scene tag; determine scene setting content based on the main scene tag and the sub-scene tag, where the scene setting content includes one or more control instructions; and control one or more smart devices based on the scene setting content to complete the scene setting.

[0080] In the embodiments of the present application, scene recognition is performed by combining the image data, audio data, and metadata of the playback content, as well as the object data of the viewing object, to obtain a main scene tag and a sub-scene tag, and scene setting content is generated based on the main scene tag and the sub-scene tag. In this way, multi-dimensional reference data is introduced in scene recognition, which can improve the intelligent dimension of scene setting. At the same time, the combined use of the main scene tag and the sub-scene tag enables the linkage of smart devices to break through the single main scene mode and introduce more specific content plot tags. In this way, the generated scene setting content can dynamically adjust multiple smart devices according to the content plot, creating an immersive atmosphere highly synchronized with the emotions of the picture, making the scene setting more rendering experience.

[0081] In some embodiments of this application, the acquisition module 501 is specifically used for: The image data, audio data, and metadata are obtained based on the target interface between the main control device and the content platform. The main control device uses its camera and / or microphone to acquire object data of the currently viewed object.

[0082] In this embodiment of the application, by combining content data and user object data, the intelligent linkage can not only respond to content but also perceive the real-time status of users, thereby improving the intelligence dimension of scene settings.

[0083] In some embodiments of this application, the processing module 502 is specifically used for: Based on a multi-feature weighted algorithm, feature extraction and feature analysis are performed on the multi-dimensional data to obtain the main scene label and the sub-scene label; And / or, Processing module 502 is specifically used for: Based on the intent recognition model, feature extraction and feature analysis are performed on the multi-dimensional data to obtain the main scene label and the sub-scene label.

[0084] In this embodiment, multi-feature weighting algorithms or intent recognition models are used to extract and analyze features from multi-dimensional data. This avoids misjudgment caused by information conflicts or redundancy during multi-dimensional data fusion. At the same time, it can dynamically evaluate the importance of each feature and accurately extract the core intent from complex or even contradictory signals. This significantly improves the accuracy and robustness of scene recognition and ensures the rationality and efficiency of the final intelligent linkage decision.

[0085] In some embodiments of this application, the processing module 502 is specifically used for: The main scene tag and the sub-scene tag are matched in the scene recommendation template library to obtain the scene setting content; and / or; Processing module 502 is specifically used for: The main scene label and the sub-scene label are processed based on the processing model to obtain the scene setting content.

[0086] In this embodiment, by introducing a template library or processing model, abstract scene labels are efficiently and standardizedly converted into specific device instructions, enabling rapid deployment, flexible expansion, and personalized customization of the linkage solution, and significantly improving the maintainability and user experience of the system.

[0087] In some embodiments of this application, the processing module 502 is further configured to: The main control device displays a first interface on its screen. This first interface is used to display the scene settings and a first interactive control. The first interactive control includes a confirmation control, a parameter adjustment control, and a rejection control. The confirmation control indicates that the scene settings are being applied. The parameter adjustment control indicates that the scene settings are being updated. The rejection control indicates that the application of the scene settings is being interrupted.

[0088] In this embodiment of the application, by providing an interactive interface that displays scene settings and interactive controls, users are given the right to know and the ability to intervene in real time regarding system commands, and personalized fine-tuning is supported, thereby avoiding the problem of users losing control due to the inability to know the specific operation in automated linkage.

[0089] In some embodiments of this application, the processing module 502 is further configured to: The main control device displays a second interface on its screen. This second interface is used to display a second interactive control, which is used to collect feedback data from the currently viewing object regarding the scene settings.

[0090] In this embodiment of the application, by providing an interactive interface to obtain user feedback information, the transparency, flexibility and user satisfaction of human-computer interaction are significantly enhanced.

[0091] In some embodiments of this application, the acquisition module 501 is further configured to: Obtain feedback information, which includes the interaction data and feedback data of the currently viewed object. The interaction data is used to characterize the interruption operation, application operation and / or parameter adjustment operation of the currently viewed object on the scene setting content. The feedback data is used to characterize the feedback of the currently viewed object on the scene setting after the scene setting content is made. Processing module 502 is also used for: Based on this feedback information, the scenario recommendation template library or processing model is updated, and this scenario recommendation template library or processing model is used to generate the scenario setting content.

[0092] In this embodiment, user actions are used as effective feedback to optimize the model, improve the intelligence of scene settings, and increase user satisfaction.

[0093] This application embodiment also provides a computer device that integrates the scene setting device of any of the smart devices provided in this application embodiment. The computer device includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor as steps in the scene setting method of the smart device in any of the embodiments described above.

[0094] This application also provides a computer device that integrates the scene setting apparatus of any of the smart devices provided in this application. For example... Figure 6 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically: The computer device may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, a power supply 603, and an input unit 604. Those skilled in the art will understand that... Figure 6 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 601 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, thereby providing overall monitoring of the computer device. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 601.

[0095] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.

[0096] The computer device also includes a power supply 603 that supplies power to the various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 603 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0097] The computer device may also include an input unit 604, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0098] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 runs the application programs stored in the memory 602 to realize various functions, as follows: Acquire multi-dimensional data, which includes image data, audio data, and metadata of the playback content currently being played by the main control device, as well as object data of the currently viewed object. The metadata is used to describe the descriptive information of the playback content. Feature extraction and feature analysis are performed on the multi-dimensional data to obtain main scene labels and sub-scene labels, wherein the sub-scene labels are sub-labels of the main scene labels; The scene setting content is determined based on the main scene label and the sub-scene label, and the scene setting content includes one or more control instructions; Based on the scenario settings, control one or more smart devices to complete the scenario setup.

[0099] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0100] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the scene setting methods for intelligent settings provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps: Acquire multi-dimensional data, which includes image data, audio data, and metadata of the playback content currently being played by the main control device, as well as object data of the currently viewed object. The metadata is used to describe the descriptive information of the playback content. Feature extraction and feature analysis are performed on the multi-dimensional data to obtain main scene labels and sub-scene labels, wherein the sub-scene labels are sub-labels of the main scene labels; The scene setting content is determined based on the main scene label and the sub-scene label, and the scene setting content includes one or more control instructions; Based on the scenario settings, control one or more smart devices to complete the scenario setup.

[0101] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0102] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0103] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0104] The above provides a detailed description of a scene setting method, apparatus, computer device, and computer-readable storage medium for a smart device according to embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for setting up scenes on a smart device, characterized in that, include: Acquire multi-dimensional data, which includes image data, audio data, and metadata of the playback content currently being played by the main control device, as well as object data of the currently viewed object. The metadata is used to describe the descriptive information of the playback content. Feature extraction and feature analysis are performed on the multi-dimensional data to obtain main scene labels and sub-scene labels, wherein the sub-scene labels are sub-labels of the main scene labels; The scene setting content is determined based on the main scene label and the sub-scene label, and the scene setting content includes one or more control instructions; Based on the scenario settings, control one or more smart devices to complete the scenario setup.

2. The method according to claim 1, characterized in that, Obtaining multi-dimensional data includes: The image data, audio data, and metadata are obtained based on the target interface between the main control device and the content platform; The object data of the currently viewed object is obtained based on the camera and / or microphone of the main control device.

3. The method according to claim 1, characterized in that, Feature extraction and feature analysis are performed on the multi-dimensional data to obtain the main scene label and sub-scene label, including: The multi-feature weighted algorithm is used to extract and analyze features from the multi-dimensional data to obtain the main scene label and the sub-scene label. And / or, Based on the intent recognition model, feature extraction and feature analysis are performed on the multi-dimensional data to obtain the main scene label and the sub-scene label.

4. The method according to claim 1, characterized in that, Determining the scene settings based on the main scene label and the sub-scene label includes: The main scene tag and the sub-scene tag are matched in the scene recommendation template library to obtain the scene setting content; and / or; The main scene label and the sub-scene label are processed based on the processing model to obtain the scene setting content.

5. The method according to any one of claims 1 to 4, characterized in that, After determining the scene settings content based on the main scene label and the sub-scene label, the method further includes: The first interface is displayed on the screen of the main control device. The first interface is used to display the scene settings content and the first interactive control. The first interactive control includes a confirmation control, a parameter adjustment control and a rejection control. The confirmation control is used to indicate the application of the scene settings content, the parameter adjustment control is used to indicate the updating of the scene settings content, and the rejection control is used to indicate the interruption of the application of the scene settings content.

6. The method according to any one of claims 1 to 4, characterized in that, After controlling one or more smart devices based on the scene settings to complete the scene settings, the method further includes: The second interface is displayed on the screen of the main control device. The second interface is used to display a second interactive control, which is used to collect feedback data from the currently viewed object on the scene settings.

7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain feedback information, which includes the interaction data and feedback data of the currently viewed object. The interaction data is used to characterize the interruption operation, application operation and / or parameter adjustment operation of the currently viewed object on the scene setting content. The feedback data is used to characterize the feedback of the currently viewed object on the scene setting after the scene setting is performed based on the scene setting content. The scene recommendation template library or the processing model is updated based on the feedback information, and the scene recommendation template library or the processing model is used to generate the scene setting content.

8. A scene setting device for a smart device, characterized in that, include: The acquisition module is used to acquire multi-dimensional data, which includes image data, audio data, and metadata of the playback content currently being played by the main control device, as well as object data of the currently viewed object. The metadata is used to describe the descriptive information of the playback content. The processing module is used to perform feature extraction and feature analysis on the multi-dimensional data to obtain main scene labels and sub-scene labels, wherein the sub-scene labels are sub-labels of the main scene labels; The scene setting content is determined based on the main scene label and the sub-scene label, and the scene setting content includes one or more control instructions; one or more smart devices are controlled based on the scene setting content to complete the scene setting.

9. A computer device, characterized in that, The computer device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 7.