Method and apparatus for interaction, and device and storage medium

By acquiring video-related data on electronic devices and matching it using a predefined model, the high latency and high cost issues of scene recognition during video playback are solved, enabling efficient and real-time scene recognition and interaction strategy determination.

WO2026067836A1PCT designated stage Publication Date: 2026-04-02BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

During video playback, existing technologies struggle to accurately identify specific scenes, resulting in poor interaction quality. Furthermore, traditional frame-skipping methods suffer from high latency and high costs.

Method used

On the electronic device, relevant video data is acquired at specified time intervals, and a predefined model is used for matching to determine the interaction strategy. This avoids frame-skipping operations on the server and reduces communication latency and computational burden.

Benefits of technology

It improves the real-time performance and interaction quality of scene recognition, reduces the computing burden and operating costs of the server, and ensures effective interaction even when offline or with an unstable network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025125768_02042026_PF_FP_ABST
    Figure CN2025125768_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a method and apparatus for interaction, and a device and a storage medium. The method comprises: during the process of playing a video, acquiring related data of the video on the basis of a specified time interval, wherein the related data comprises audio-visual data of the video and data detected by a target device that plays the video; on the basis of a predetermined model, executing matching on the related data and at least one predetermined scene, so as to obtain a matching result; and in response to the matching result indicating that the related data matches the at least one predetermined scene, determining an interaction strategy corresponding to the at least one predetermined scene.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, apparatus and storage medium for interaction

[0001] The present application claims priority to the Chinese patent application No. 202411391769.0, filed on September 30, 2024, entitled “Method, device, apparatus and storage medium for interaction”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to a method, device, apparatus and storage medium for interaction. BACKGROUND

[0003] In the current mobile device video playing scene, it is of great significance to accurately identify a specific scene in the video playing process. For the identification of a specific scene, usually depends on the frame extraction of the video. But the frame extraction method is easy to fail to accurately determine the interaction strategy corresponding to the scene due to missing frames. Thus leading to poor overall interaction quality. SUMMARY

[0004] In a first aspect of the present disclosure, a method for interaction is provided. The method can include: in a video playing process, obtaining related data of the video based on a specified time interval. Based on a predetermined model, performing matching of the related data with at least one predetermined scene to obtain a matching result. In response to the matching result indicating that the related data matches the at least one predetermined scene, determining an interaction strategy corresponding to the at least one predetermined scene.

[0005] In a second aspect of the present disclosure, a device for interaction is provided. The device can include: a related data obtaining module configured to, in a video playing process, obtain related data of the video based on a specified time interval. A matching result determining module configured to, based on a predetermined model, perform matching of the related data with at least one predetermined scene to obtain a matching result. An interaction strategy determining module configured to, in response to the matching result indicating that the related data matches the at least one predetermined scene, determine an interaction strategy corresponding to the at least one predetermined scene.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the electronic device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The medium has stored thereon a computer program, which, when executed by a processor, implements the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the method provided in various optional manners in the first aspect of the present application. In other words, the computer instructions are executed by the processor to implement the method provided in various optional manners in the first aspect of the present application.

[0009] It should be understood that the content described in this section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:

[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG. 2 shows a flowchart of a method for interaction according to some embodiments of the present disclosure;

[0013] FIG. 3 shows an example diagram of interaction of a server and an electronic device according to some embodiments of the present disclosure;

[0014] FIG. 4 shows a schematic diagram of interaction of an electronic device end according to some embodiments of the present disclosure;

[0015] FIG. 5 shows a schematic structural block diagram of an interaction device according to some embodiments of the present disclosure; and

[0016] FIG. 6 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail by referring to the drawings. Although certain embodiments of the present disclosure are illustrated in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, these embodiments are provided so as to more thoroughly and completely understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] In the description of embodiments of the disclosure, the term "comprising" and similar terms are to be interpreted as open-ended, i.e., "including but not limited to". The term "based on" is to be interpreted as "based, at least in part, on". The term "one embodiment" or "the embodiment" is to be interpreted as "at least one embodiment". The term "some embodiments" is to be interpreted as "at least some embodiments". Other explicit and implicit definitions can also be included below.

[0019] In this document, unless explicitly stated, performing a step "in response to A" does not mean performing the step immediately after A, but can include one or more intermediate steps.

[0020] It can be understood that the data involved in the technical solutions of the present application (including but not limited to the data itself, obtaining, using, storing or deleting) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained by appropriate means according to relevant laws and regulations, wherein the relevant user can include any type of right subject, such as individual, enterprise, group.

[0022] For example, in response to receiving the active request of the user, the prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require obtaining and using the information of the relevant user, so that the relevant user can voluntarily choose whether to provide information to the software or hardware such as electronic device, application program, server or storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.

[0023] As an optional but non-limiting implementation manner, in response to receiving the active request of the relevant user, the prompt information is sent to the relevant user, for example, in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.

[0024] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0025] FIG. 1 illustrates a schematic diagram of an environment 100 in which embodiments of the present disclosure can be implemented. An interactive application scenario is illustrated in the environment 100 of FIG. 1. In the environment of FIG. 1, an electronic device 110 can present a page 150 of a target application 120. The page 150 can include various types of pages 150 that can be provided by the application 120. In an embodiment, the page can be a page related to video playing. For example, a game page, a live page, etc. The electronic device 110 can perform recognition on the presented page related to video playing to determine whether the page can match a predetermined scene 160. If at least one predetermined scene is matched, an interaction policy 170 corresponding to the predetermined scene can be determined. During the running of the target application 120, the electronic device 110 can invoke one or more models, e.g., capabilities of the models, to determine whether the page can match the predetermined scene.

[0026] In some embodiments, the electronic device 110 communicates with a server 130 to implement provisioning of services to the electronic device 110. Illustratively, the provisioning of services to the electronic device 110 can include provisioning of models. The provisioning can include providing download of models or update of models, etc. The electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media player, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combinations of the aforementioned and the like, including accessories and peripherals for these devices, or any combinations thereof. In some embodiments, the electronic device 110 can also be capable of supporting any type of interface to a user (such as "wearable" circuitry, etc.). The server 130 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, etc.

[0027] It should be understood that the structures and functions of the various elements in the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.

[0028] For a traditional determination scheme of a scene corresponding to a video, usually a machine learning model deployed on a server side performs a timing frame extraction on a page of a target application presented by an electronic device, so as to complete the scene corresponding to the video through recognition of the frame extraction result. However, this method has problems of high delay and high cost, especially the communication delay between the server and the client, which can cause insufficient real-time performance of the recognition. With the increase of the number of clients, the operation burden and cost of the server also increase. Therefore, the traditional method cannot solve the problems of high delay and high cost, resulting in a poor interaction quality presented on the electronic device.

[0029] To solve the above problems, in an embodiment of the present disclosure, a method for interaction is proposed. In the electronic device 110, during the video playing process, relevant data of the video is obtained based on a specified time interval, the relevant data including audio-visual data of the video and data detected by a target device playing the video. Based on a predetermined model, the relevant data is matched with at least one predetermined scene to obtain a matching result. In response to the matching result indicating a successful match with the at least one predetermined scene, an interaction strategy corresponding to the at least one predetermined scene is determined.

[0030] Through the above process, the frame extraction operation of the server 130 on the electronic device 110 is avoided. Instead of the frame extraction operation, the relevant data of the video is obtained in the electronic device 110 based on a specified time interval, and the matching with the predetermined scene is completed based on the relevant data. Through real-time processing and recognition of the relevant data, there is no need to upload the data to the server for analysis. This method eliminates the communication delay between the electronic device 110 and the server 130, significantly improves the real-time performance of the scene recognition. In addition, the server 130 does not need to perform frequent frame extraction and complex calculation, reducing the calculation burden and operation cost of the server 130. In actual application, the electronic device 110 can utilize the local processing capability to perform the matching operation of the scene recognition and the interaction strategy. This not only improves the speed and accuracy of the recognition, but also allows the electronic device 100 to still perform the scene recognition and the execution of the interaction strategy in an offline or unstable network condition. Therefore, the interaction quality and efficiency can be improved.

[0031] FIG. 2 shows an example flow 200 of a method for interaction according to some embodiments of the present disclosure. For ease of discussion, the flow 200 will be described with reference to the environment of FIG. 1. In the application environment 110, different pages of the target application 120 presented by the electronic device 110 can be determined to correspond to predetermined scenes 160. Based on the determined predetermined scenes 160, an interaction strategy 170 matching the predetermined scenes 160 can be executed. In an embodiment of the present disclosure, different pages can be, for example, game pages, live broadcast pages, etc.

[0032] At block 201, the electronic device 110 acquires relevant data of a video based on a specified time interval during video playback. The video being played can be a game video, a live video, etc. In this disclosure, the entire process is described in detail taking the live game video as an example.

[0033] During the process of playing the video, the electronic device 110 can acquire audio data and picture data of the video in real time based on a specified time interval. The audio data can include sound effects, background music, voice, etc. The picture data can include video image frames. These data are received through the audio and video receiving module built in the electronic device 110 and stored in the cache of the device for subsequent processing. The audio data and picture data are hereinafter referred to as audio and picture data.

[0034] In addition, the relevant data of the video can include data detected by the target device playing the video, i.e. data detected by the target device playing the video, which can also be referred to as device-related data. Specifically, during the process of playing the video, the electronic device 110 can also detect data related to the video. Exemplarily, it can include data provided by sensors such as accelerometer, gyroscope, GPS, etc. in the electronic device 110. These data can reflect the motion state, position, etc. of the electronic device 110 during the video playback process. It can also include the type of current network connection of the electronic device 110 (such as Wi-Fi, 4G / 5G), signal strength, delay (ping value), etc. In addition, it can also include CPU usage, memory usage, battery status, etc.

[0035] During the process of playing the video, the electronic device 110 periodically receives the relevant data of the video at a specified time interval (e.g. once per second). The specified time interval can be dynamically adjusted according to the specific application requirements to achieve the best balance point, which ensures the real-time nature of the relevant data acquisition and does not consume the system resources of the electronic device 110 excessively. The dynamic adjustment of the time interval will be described in detail later.

[0036] At block 202, the electronic device 110 matches the relevant data with at least one predetermined scene based on a predetermined model to obtain a matching result. The electronic device 110 has a predetermined model built in or downloaded. The model is pre-trained to identify different predetermined scenes. Taking the video as a game live video as an example, the predetermined scenes can include game intermission scene, game highlight scene, game result summary scene, game data summary scene, game network status scene, etc.

[0037] During the video playback, the electronic device 110 can input the relevant data received at specified time intervals into the predetermined model. The relevant data is analyzed and processed by the predetermined model to evaluate the relevance (e.g., matching degree) of the relevant data to each predetermined scene. The relevance can be used as the matching result of the relevant data to each predetermined scene.

[0038] At block 203, the electronic device 110 determines an interaction strategy corresponding to the at least one predetermined scene in response to the matching result indicating that the relevant data matches the at least one predetermined scene.

[0039] The electronic device 110 calculates the matching result, i.e., the size of the relevance, by the predetermined model. For each scene, a matching result threshold can be set. For example, the matching result threshold is set to 90%. If the relevance is not lower than the matching result threshold, it indicates that the relevant data obtained at the moment can match the corresponding predetermined scene. Conversely, if the relevance is lower than the matching result threshold, it indicates that the relevant data obtained at the moment fails to match the corresponding predetermined scene.

[0040] Based on the identified predetermined scene, the electronic device 110 can select an interaction strategy corresponding to the scene. Each predetermined scene has a set of predefined interaction strategies. For example, for the game interlude scene, the corresponding interaction strategy can include inserting content that enhances the interactive atmosphere, such as displaying game-related information, player interaction information, etc. For the game highlight moment scene, the corresponding interaction strategy can include capturing the current highlight frame and using it as a live broadcast cover, or marking it as a highlight segment for user review. For the game result summary scene, the corresponding interaction strategy can include sending a guessing reward, displaying game data statistics, or pushing a summary report to the player. For the game data statistics scene, the statistics can be displayed on the live broadcast frame as auxiliary information. For the game network status scene, for example, when the network condition is detected to be poor (e.g., ping value is too high), the electronic device 110 can automatically adjust its own computing resource occupation to ensure smooth video playback.

[0041] Through the above process, the relevant data acquisition of the video is performed at the electronic device 110, and the scene recognition is performed based on the relevant data of the video, which significantly reduces the delay and improves the real-time performance. The above process weakens the interaction between the electronic device 110 and the server 130, thus reducing the computational burden and communication cost of the server 130 and reducing the overall operating cost.

[0042] In some embodiments of the present disclosure, the relevant data includes audio-visual data of the video and data detected by the target device playing the video. The specific process of matching the relevant data with the at least one predetermined scene can include: combining the relevant data and the identification corresponding to the at least one predetermined scene into input information as the input of the predetermined model. By processing the input information by using the predetermined model, a correlation degree is obtained, which indicates the matching degree of the relevant data with each predetermined scene in the at least one predetermined scene. The matching result is obtained based on the correlation degree.

[0043] Each predetermined scene can correspond to an identification. The electronic device 110 combines the identification of the at least one predetermined scene to be identified and the obtained relevant data to form complete input information. For example, if the game intermission scene and the game highlight moment scene are to be identified, the identifications of the two scenes will be added to the input information. If only the game highlight moment scene is to be identified, only the identification corresponding to the game highlight moment scene will be added to the input information.

[0044] The input information of the predetermined model includes two parts: the relevant data and the identification of the predetermined scene. The relevant data is the various data received from the video playing process, and the predetermined scene identification is the predefined scene type. After the data at the current time is received, the electronic device 110 integrates these input data in the form of key-value pairs as the input information of the predetermined model. In addition, the identification of the scene to be identified is added to the input information. Exemplarily, the input information of the predetermined model can include the following contents:

[0045] Identified scene: [“game intermission scene”, “game highlight moment scene”, “game result summary scene”];

[0046] Microphone data: [0, 1, 0, 0, 3,...];

[0047] Frame picture data: [255, 255, 0, 0, 255,...];

[0048] Sensor data (x, y, z): [111, 200, 300].

[0049] The microphone data can correspond to the volume received by each sound receiving point. The frame picture data can correspond to the color representation of each pixel point of the image. The sensor data can correspond to the coordinate representation of the electronic device 110 on the x-axis, y-axis and z-axis. Based on the input information, the predetermined model can obtain a correlation degree result. The correlation degree result is an indication value indicating the matching degree of the relevant data with each predetermined scene. For example, a correlation degree of 90% indicates that the relevant data is highly matched with a predetermined scene, and 10% indicates a lower matching degree.

[0050] The electronic device 110 obtains a relevance result of the relevant data corresponding to each predetermined scene based on the output of the predetermined model. The electronic device 110 determines a matching result based on the relevance of each predetermined scene. The matching result is determined according to the relevance result. The higher the relevance result, the higher the matching degree of the relevant data and the predetermined scene.

[0051] Through the above steps, the electronic device 110 can effectively use the predetermined model to process the relevant data in the video playing process, calculate the matching degree with the predetermined scene, and determine the matching result based on the relevance. This process can accurately identify different scenes and execute corresponding interaction strategies.

[0052] In some embodiments of the present disclosure, the specified time interval can be adjusted based on the matching result. Specifically, a target relevance is selected from the relevance of the relevant data and each predetermined scene. Based on the target relevance, an adjustment coefficient is determined, and the adjustment coefficient is inversely proportional to the target relevance. Based on the adjustment coefficient, the specified time interval is adjusted, and the adjusted specified time interval is proportional to the adjustment coefficient.

[0053] The target relevance in the relevance of the relevant data and each predetermined scene can be the maximum value in each relevance. For example, the relevance of the relevant data and the game interlude scene, the game highlight moment scene, and the game result summary scene is 18%, 50%, and 90%, respectively. Then 90% can be selected as the target relevance.

[0054] Based on the target relevance, the adjustment coefficient x can be determined. The adjustment coefficient x can be expressed as expression (1):

[0055] x = 1 - confidence (1)

[0056] In expression (1), confidence can be the target relevance, that is, the adjustment coefficient is inversely proportional to the target relevance.

[0057] Based on the adjustment coefficient, the specified time interval can be adjusted. The adjustment of the specified time interval can be expressed as expression (2): interval = min_interval + (max_interval - min_interval) * (1 - onfidence) (2)

[0058] In expression (2), min_interval and max_interval correspond to the minimum time interval and the maximum time interval, respectively. The minimum time interval and the maximum time interval can be set based on the model, performance, and other parameters of the electronic device 110. For example, the minimum time interval can be set to 1 second, and the maximum time interval can be set to 10 seconds. Based on the adjustment coefficient, the specified time interval is adjusted, and the adjusted specified time interval is proportional to the adjustment coefficient.

[0059] For example, the first time the relevant data of the video is acquired and matched with at least one predetermined scene execution, the maximum correlation determined is 0.99. Then, based on the minimum time interval of 1 second and the maximum time interval of 10 seconds, it can be determined that the time interval between the second acquisition and the first acquisition is 1.09 seconds. The second time the relevant data of the video is acquired and matched with at least one predetermined scene execution, if the maximum correlation determined is 0.01. Then, based on the minimum time interval of 1 second and the maximum time interval of 10 seconds, it can be determined that the time interval between the third acquisition and the second acquisition is 9.91 seconds.

[0060] Through the above process, the electronic device 110 can dynamically adjust the time interval based on the matching result each time to achieve a more flexible and efficient scene recognition and processing mechanism. This way of dynamically adjusting the specified time interval significantly improves the response speed and resource utilization efficiency of the electronic device 110, solving the high delay and high cost problem in the traditional method.

[0061] The working principle of the predetermined model is introduced above. Next, the acquisition method and the configuration method of the predetermined model are introduced. In some embodiments of the present disclosure, the electronic device 110 sends an acquisition request for acquiring model release information in response to a trigger instruction. In response to the model release information indicating that the predetermined model needs to be acquired, the predetermined model is acquired based on the model release information.

[0062] The trigger instruction can be in response to the start of the target application 120. Alternatively, the trigger instruction can be in response to a received instruction of the user, etc. Based on the trigger instruction, the electronic device 110 sends an acquisition request for acquiring model release information to the server 130 to acquire the preset model release information. The server 130 provides the model release information in response to the acquisition request. Illustratively, the model release information includes the version of the currently released preset model, the download address, and the like.

[0063] After the electronic device 110 receives the model release information, it is determined whether the version of the newly released preset model is consistent with the version of the existing preset model in the electronic device 110. If the version of the newly released model is inconsistent with the existing version, the electronic device 110 will acquire the latest predetermined model based on the download address in the model release information.

[0064] Through the above process, the electronic device 110 can download and update the preset model when a new model is detected to be released. This function ensures that the electronic device 110 always runs the latest and optimized model, improving the accuracy and effectiveness of scene recognition and interaction strategies.

[0065] After obtaining the latest version of the preset model, and before obtaining the related data, the electronic device 110 first obtains the configuration information corresponding to the predetermined model. If the configuration information indicates that the use state of the predetermined model is available, the predetermined model can be loaded and the related data can be obtained.

[0066] Through the above process, after obtaining the configuration information, the electronic device 110 obtains the related data in response to the configuration information indicating that the use state of the predetermined model is available. This step ensures that data processing and scene recognition are only performed when the preset model meets the requirements and is applicable.

[0067] Based on the matching result determined by the preset model that meets the requirements and is applicable, the interaction strategy configuration information indicating the correspondence between the target scene in the at least one predetermined scene and the target interaction strategy can be obtained. Based on the interaction strategy configuration information, the interaction strategy corresponding to the at least one predetermined scene is determined.

[0068] The interaction strategy configuration information indicates the correspondence between the target scene in the predetermined scene and the target interaction strategy. Based on the obtained interaction strategy configuration information, the electronic device 110 can accurately determine the interaction strategy corresponding to each predetermined scene. For example, providing interactive content during game breaks, automatically capturing highlight scenes during game highlights, etc. This interaction based on strategy configuration information enables the system to provide the most suitable interaction strategy in different use scenarios, improving interaction efficiency. By dynamically adjusting the interaction strategy, the system can also optimize resource utilization, avoiding unnecessary computation and data transmission.

[0069] For the predetermined model, it can be converted to a format supported by the target device during its release process. The server 130 will convert the predetermined model to a format supported by the target device during the release stage of the predetermined model. The reason for format conversion is that the predetermined model developed on the server 130 side usually has strong computing power and different operating systems, while the electronic device 110 as a mobile terminal usually has limited computing resources and different operating systems. In order to adapt to these differences, the predetermined model needs to be format-converted on the server 130 side so that the format-converted predetermined model can run efficiently on the electronic device 110.

[0070] The electronic device 110 can determine the model publishing information corresponding to the predetermined model in the process of obtaining the predetermined model, and the model publishing information includes at least one of version information and a download link of the predetermined model. Based on the model publishing information, the predetermined model is downloaded to the target device.

[0071] The server 130 performs registration on the predetermined model in the publishing stage of the predetermined model. The registered predetermined model generates version information, a download link, and the like. The version information, the download link, and the like are encapsulated as model publishing information. After obtaining the model publishing information, the electronic device 110 can determine whether to perform download on the predetermined model based on the model publishing information. For example, if the newly published predetermined model is the same as the predetermined model version already existing in the electronic device 110, there is no need to repeat the download. Only when the newly published predetermined model is different from the predetermined model version already existing in the electronic device 110, the electronic device 110 needs to update the predetermined model by downloading. If the download is performed, the download can be completed based on the download link.

[0072] FIG. 3 shows an interaction diagram of the server 130 and the electronic device 110 according to some embodiments of the present disclosure. On the server 130 side, the predetermined model generation and delivery process includes the following steps: performing predetermined model development 311 in block 311. According to actual needs such as performance optimization and function expansion, a new predetermined model is developed. Perform predetermined model training in block 312. Perform predetermined model format conversion in block 313. Because the environment of the server 130 side and the electronic device 110 side is different, the predetermined model produced by the server 130 needs to be converted into a format supported by the electronic device 110. Perform predetermined model deployment in block 314. After the predetermined model is ready, the registration and deployment of the model can be performed to generate corresponding model publishing information. Perform predetermined model delivery in block 315. Generate configuration information, and based on the request of the electronic device 110 side, deliver the model to the electronic device 110. This process ensures that the latest predetermined model can be dynamically obtained and used during the execution of the application program, and realizes real-time updating and dynamic adjustment.

[0073] At the electronic device 110 end, the predetermined model acquisition and the process includes the following steps: taking the target application start as the trigger condition. In block 321, the judgment of the predetermined model version is performed. The electronic device 110 will judge whether the existing predetermined model version is consistent with the latest published predetermined model version in response to the acquired model publishing information. If the version inconsistency is detected, the latest version of the predetermined model needs to be downloaded. In block 322, the predetermined model download is performed. The electronic device 110 downloads the latest predetermined model from the server 130 to the local storage. In block 323, the predetermined model loading is performed. After the download is completed, the electronic device 110 judges the availability of the predetermined model based on the configuration information. If available, the predetermined model is preloaded. In block 324, the predetermined model running is performed. During the video playing process, the electronic device 110 acquires the relevant data of the video based on the specified time interval, including the audio-visual data and the data detected by the target device. These data are processed by the predetermined model, matched with the predetermined scene execution, the matching result is obtained, and the corresponding interaction strategy is determined. These steps ensure that the electronic device can dynamically acquire the latest predetermined model and efficiently run in actual application, and realize the dynamic adjustment of scene recognition and interaction strategy.

[0074] FIG. 4 shows an interaction schematic diagram of the electronic device end according to some embodiments of the present disclosure. In block 411, the electronic device 110 takes the target application start as the trigger instruction. In block 412, the model publishing information is acquired. In block 413, based on the model publishing information, it is judged whether the model has been updated. If it has been updated, in block 414, the predetermined model is updated, and in block 415, the predetermined model is preloaded. If it has not been updated, the existing predetermined model is directly loaded in block 415.

[0075] After the predetermined model is loaded, in block 416, the video is played. The electronic device 110 will first judge whether the predetermined model is available. If not available, the whole process ends. If available, in block 418, the relevant data of the video is acquired, such as the audio-visual data of the video and the data detected by the target device playing the video, etc. In block 419, based on the predetermined model, the relevant data is matched with at least one predetermined scene execution, and the matching result is obtained. In block 420, based on the matching result, it is determined whether the relevant data matches at least one predetermined scene. In block 421, based on the matching result, the specified time interval is dynamically adjusted. For example, if the matching result indicates that the relevance to at least one predetermined scene is high, the receiving time interval can be shortened. Conversely, if the relevance to each predetermined scene is low, the receiving time interval can be lengthened. Exemplarily, the high or low of the relevance can be judged according to the threshold. In block 422, the interaction strategy corresponding to at least one predetermined scene is determined.

[0076] FIG. 5 shows a schematic structural block diagram of an apparatus 500 for interaction, according to some embodiments of the present disclosure. The apparatus 500 may, for example, be implemented in or included in the model training system and / or the electronic device 110. Various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0077] As shown, the apparatus 500 includes a related data obtaining module 501 configured to obtain, during a video playing process, related data of a video based on a specified time interval, the related data including audio-visual data of the video and data detected by a target device playing the video. A matching result determining module 502 is configured to perform matching of the related data with at least one predetermined scene based on a predetermined model, to obtain a matching result. An interaction strategy determining module 503 is configured to, in response to the matching result indicating that the related data matches the at least one predetermined scene, determine an interaction strategy corresponding to the at least one predetermined scene.

[0078] In some embodiments of the present disclosure, the matching result determining module 502 can be specifically configured to: compose the related data and an identifier corresponding to the at least one predetermined scene into input information, as input of the predetermined model. By processing the input information by using the predetermined model, a correlation degree is obtained, the correlation degree indicating a matching degree of the related data with each of the at least one predetermined scene. The matching result is obtained based on the correlation degree.

[0079] In some embodiments of the present disclosure, the apparatus 500 further includes a specified time interval adjusting module configured to: adjust the specified time interval based on the matching result.

[0080] In some embodiments of the present disclosure, the specified time interval adjusting module can be specifically configured to: select a target correlation degree from the correlation degrees of the related data and the respective predetermined scenes. Based on the target correlation degree, an adjustment coefficient is determined, the adjustment coefficient being inversely proportional to the target correlation degree. Based on the adjustment coefficient, the specified time interval is adjusted, the adjusted specified time interval being proportional to the adjustment coefficient.

[0081] In some embodiments of the present disclosure, the apparatus 500 further includes a predetermined model obtaining module. The module is configured to: in response to a trigger instruction, send an obtaining request for obtaining model publishing information. In response to the obtained model publishing information indicating that the predetermined model needs to be obtained, the predetermined model is obtained based on the model publishing information.

[0082] In some embodiments of the present disclosure, the apparatus 500 further includes a detection module configured to: before obtaining the related data, obtain configuration information corresponding to the predetermined model. In response to the configuration information indicating that a use state of the predetermined model is available, the related data is obtained.

[0083] In some embodiments of the present disclosure, the interaction policy determination module 503 is specifically configured to: obtain interaction policy configuration information, the interaction policy configuration information indicating a correspondence between a target scene in at least one predetermined scene and a target interaction policy. Based on the interaction policy configuration information, determine the interaction policy corresponding to the at least one predetermined scene.

[0084] In some embodiments of the present disclosure, the predetermined model is converted to a format supported by the target device, and the apparatus 500 is further configured to download the predetermined model to the target device.

[0085] In some embodiments of the present disclosure, downloading the predetermined model to the target device includes: determining model release information corresponding to the predetermined model, the model release information including at least one of version information and a download link of the predetermined model. Based on the model release information, downloading the predetermined model to the target device.

[0086] In some embodiments of the present disclosure, downloading the predetermined model to the target device includes: determining configuration information corresponding to the predetermined model, the configuration information indicating a target device to which the predetermined model is adapted. In response to determining that the configuration information indicates that the target device is adapted to the predetermined model, downloading the predetermined model to the target device.

[0087] FIG. 6 illustrates a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented. It should be understood that the electronic device 600 illustrated in FIG. 6 is merely exemplary and should not be construed as any limitation of the functionality and scope of the embodiments described herein. The electronic device 600 illustrated in FIG. 6 can include or be implemented as the electronic device 110 and / or the model application system of FIG. 1, or the apparatus 500 of FIG. 5.

[0088] As shown in FIG. 6, the electronic device 600 is in the form of a general electronic device. The components of the electronic device 600 can include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 can be an actual or virtual processor and can perform various processing according to programs stored in the memory 620. In a multi-processor system, multiple processors perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0089] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by the electronic device 600 and includes both volatile and nonvolatile media, removable and non-removable media. The memory 620 can be volatile (such as register, cache, RAM), non-volatile (such as ROM, EEPROM, flash memory), or some combination of the two. The storage device 630 can be a removable or non-removable media, and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 600.

[0090] The electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM) can be provided. In such instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0091] The communication unit 640 enables communications with other electronic devices over a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with one another over a communication connection. As such, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0092] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 640, as needed, one or more devices that enable a user to interact with the electronic device 600, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 600 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0093] According to an example implementation of the present disclosure, a computer readable storage medium is provided, having computer executable instructions stored thereon, wherein the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method described above.

[0094] According to an example implementation of the present disclosure, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional manners in FIGS. 2 to 4, and therefore, details will not be repeated here.

[0095] Various aspects of the disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and computer program products according to implementations of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0096] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can comprise a non-transitory computer readable medium. The instructions stored in the computer readable storage medium, which can comprise a non-transitory computer readable medium, can cause a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable medium having instructions stored therein comprises an article of manufacture including a manufacture that implements the function / act specified in the flowchart and / or block diagram block or blocks.

[0097] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0098] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk, or memory stick, can also be used to implement the present disclosure. The computer program product of the present disclosure can also be provided as a service to download and use the computer program over a network, such as the Internet.

[0099] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and substitutions are possible without departing from the scope and spirit of the described implementations. Many modifications and variations of the present disclosure are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims, the present disclosure can be practiced otherwise than as specifically described. While the present disclosure has been described with reference to the implementation figures, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt to a particular situation and the teachings of the present disclosure to a specific implementation, without departing from the central novel teachings of the application. The implementation(s) illustrated and described herein are meant only to serve as examples. Departures in form and detail are within the scope of the disclosure. Therefore, one skilled in the art can restructure the implementation(s) as needed, while still adhering to the principles of the present disclosure.

Claims

1.A method for interaction, comprising: acquiring, during a video playing process, related data of the video based on a specified time interval; performing matching between the related data and at least one predetermined scene based on a predetermined model to obtain a matching result; and determining an interaction strategy corresponding to the at least one predetermined scene in response to the matching result indicating that the related data matches the at least one predetermined scene. 2.The method of claim 1, wherein the related data comprises audio-visual data of the video and data detected by a target device playing the video, and obtaining the matching result comprises: forming identification corresponding to the related data and the at least one predetermined scene into input information as input of the predetermined model; obtaining a correlation degree by processing the input information by using the predetermined model, the correlation degree indicating a matching degree of the related data and each of the at least one predetermined scene; and obtaining the matching result based on the correlation degree. 3.The method of claim 1, further comprising: adjusting the specified time interval based on the matching result. 4.The method of claim 3, wherein the matching result corresponds to a correlation degree of the related data and each of the at least one predetermined scene, and adjusting the specified time interval based on the matching result comprises: selecting a target correlation degree from the correlation degrees of the related data and each of the at least one predetermined scene; determining an adjustment coefficient based on the target correlation degree, the adjustment coefficient being inversely proportional to the target correlation degree; and adjusting the specified time interval based on the adjustment coefficient, the adjusted specified time interval being proportional to the adjustment coefficient. 5.The method of claim 1, further comprising: sending an acquisition request for acquiring model publishing information in response to a trigger instruction; and acquiring the predetermined model based on the model publishing information in response to the acquired model publishing information indicating that the predetermined model needs to be acquired. 6.The method of claim 1, further comprising: acquiring configuration information corresponding to the predetermined model before acquiring the related data; and acquiring the related data in response to the configuration information indicating that a use state of the predetermined model is available. 7.The method of claim 1, wherein determining the interaction strategy corresponding to the at least one predetermined scene comprises: acquiring interaction strategy configuration information, the interaction strategy configuration information indicating a corresponding relationship between a target scene of the at least one predetermined scene and a target interaction strategy; and determining the interaction strategy corresponding to the at least one predetermined scene based on the interaction strategy configuration information. downloading the predetermined model to the target device. 9.An apparatus for interaction, comprising: a related data acquisition module configured to acquire, during a video playing process, related data of the video based on a specified time interval; a matching result determination module configured to perform matching between the related data and at least one predetermined scene based on a predetermined model to obtain a matching result; and a matching result determination module configured to perform matching between the related data and at least one predetermined scene based on a predetermined model to obtain a matching result; and ​ 8. The method of claim 1, wherein the predetermined model is converted to a format supported by the target device, and the method further comprises: ​ ​ ​ ​ ​ An interaction policy determining module, configured to determine an interaction policy corresponding to the at least one predetermined scene in response to the matching result indicating that the related data matches the at least one predetermined scene. 10.An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1 to 8. 11.A computer readable storage medium having computer executable instructions stored thereon, the computer executable instructions executable by a processor to implement the method according to any one of claims 1 to 8. 12.A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Internet addiction detection device and method based on user-computer interactive events

    CN103413054A

  • Information display method and device, electronic equipment and storage medium

    CN111694983A

  • Intelligent method for high-risk operation live broadcast interruption based on scene recognition

    CN111818356A

  • Live broadcast interaction method and device, equipment and storage medium

    CN117640982A

  • Scene recognizing device

    JP2000293694A