A method and device for controlling the usage time of electronic devices based on deep learning

Through multi-level deep learning models and an adaptive edge inference framework, the problems of imprecise screen content recognition, high resource consumption, and privacy risks in existing technologies have been solved, and refined recognition and personalized control of screen content have been achieved, thereby improving the user experience.

CN120448237BActive Publication Date: 2025-09-12SENARY TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510953672.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-12
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing technologies are unable to perform fine-grained recognition of screen content, consume a lot of resources, pose privacy risks, have a single management and control strategy, provide poor user experience, and easily induce evasive behavior.

Method used

It adopts a multi-level deep learning model and an adaptive edge inference framework, and generates personalized usage time management strategies by performing multi-dimensional feature extraction and pattern recognition on screen image sequences. The calculation is mainly completed locally, reducing resource consumption and providing privacy protection.

Benefits of technology

It achieves refined recognition of screen content, distinguishes different usage modes within the same application, reduces computing resource consumption, extends battery life, generates differentiated management and control strategies, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448237B_ABST
    Figure CN120448237B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for controlling the usage time of electronic devices based on deep learning, which obtains the user interface data of the current screen image sequence and the current device related information, and determines the current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship; determines the corresponding mode weight and the basic threshold value of the usage time of the electronic device based on the current scene label; determines the working period, continuous use time and actual use time of the electronic device under the current scene label based on the current device related information; generates a time management strategy for the electronic device based on the mode weight, basic threshold value of usage time, working period, continuous use time and actual use time. Through multi-level deep learning, the screen content can be finely identified and different usage modes within the same application can be distinguished; through multi-dimensional factors and usage scenarios, differentiated management strategies can be generated to improve user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic control, and in particular to a method and device for controlling the usage time of electronic devices based on deep learning. Background Art

[0002] With the rapid development of modern technology, electronic devices (such as smartphones, tablets, laptops, etc.) are increasingly used in daily life. Their convenience and versatility are deeply loved by users of all ages. However, screen time management has become an important issue, especially for adolescent users who have relatively weak self-control and are in a critical period of growth. This issue is particularly prominent and important.

[0003] Screen time management refers to a systematic approach that uses technology to monitor, control, and optimize the amount of time users spend using electronic devices. Its core goal is to prevent the negative impacts of excessive screen time on health (such as myopia and sleep disorders), attention span, and life balance, while ensuring users' normal use needs.

[0004] Existing technologies mainly perform time control based on application package names or simple rules. They are unable to identify screen content in a refined manner, consume a lot of resources, pose privacy risks, and have a single control strategy, resulting in a poor user experience and easily leading to evasive behavior. Summary of the Invention

[0005] In view of the above problems, the present application is proposed to provide a method and device for controlling the usage time of an electronic device based on deep learning to overcome the above problems or at least partially solve the above problems, including:

[0006] A method for controlling the usage duration of electronic devices based on deep learning, the method being applied to establishing a correspondence between user interface data and device-related information of a device screen image sequence and corresponding scene labels through an artificial intelligence model; wherein the user interface data includes a user operation interface layout and visual elements of the user operation interface; the user operation interface layout includes the spatial arrangement, proportional relationship, and functional division of each component of the user operation interface; visual elements refer to the visual components that do not change over time that constitute a video image; and the device-related information includes operation time information and historical scene recognition information.

[0007] The method comprises:

[0008] Acquire user interface data of a current screen image sequence and current device related information, and determine a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship;

[0009] Determining a corresponding mode weight and a basic threshold of the usage time of the electronic device according to the current scene label;

[0010] Determine the working period, continuous usage time and actual usage time of the electronic device under the current scene tag based on the current device related information;

[0011] A duration management strategy for the electronic device is generated according to the mode weight, the basic usage duration threshold, the working period, the continuous usage duration, and the actual usage duration.

[0012] Furthermore, the step of determining a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship includes:

[0013] Determine spatial characteristics and temporal characteristics of the user interface data and the current device related information;

[0014] generating a fusion feature based on the spatial feature and the temporal feature;

[0015] The current scene label is generated according to the fusion feature.

[0016] Furthermore, the step of determining a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship includes:

[0017] Determine spatial characteristics and temporal characteristics of the user interface data and the current device related information;

[0018] generating a fusion feature based on the spatial feature and the temporal feature;

[0019] determining enhanced features of the current device-related information;

[0020] The current scene label is generated according to the fused feature and the enhanced feature.

[0021] Furthermore, the step of generating a duration management policy for the electronic device based on the mode weight, the basic usage duration threshold, the working period, the continuous usage duration, and the actual usage duration includes:

[0022] Generate a corresponding dynamic usage time threshold according to the basic usage time threshold, the working period, and the continuous usage time;

[0023] generating a weighted usage duration according to the mode weight and the actual usage duration;

[0024] A duration management strategy for the electronic device is generated according to the duration dynamic threshold and the weighted usage duration.

[0025] Furthermore, the step of generating a corresponding dynamic usage duration threshold based on the usage duration basic threshold, the working period, and the continuous usage duration includes:

[0026] Determining a time period coefficient corresponding to the current scene label according to the working time period;

[0027] determining a fatigue factor based on the continuous use duration;

[0028] The dynamic usage duration threshold is generated according to the basic usage duration threshold, the time period coefficient and the fatigue coefficient.

[0029] Furthermore, the step of generating a duration management policy for the electronic device based on the duration dynamic threshold and the weighted usage duration includes:

[0030] Obtaining a user profile, and determining the intervention sensitivity of the corresponding user based on the user profile;

[0031] generating a remaining usage time according to the dynamic usage threshold and the weighted usage time;

[0032] Determine the warning level based on the remaining usage time;

[0033] An intervention strategy corresponding to the user is generated according to the intervention sensitivity and the warning level.

[0034] Furthermore, it also includes:

[0035] Obtaining user feedback on the time management strategy;

[0036] The duration control strategy is optimized based on the feedback information.

[0037] A deep learning-based device for controlling the usage duration of an electronic device, the device being configured to establish a correspondence between user interface data and device-related information of a device screen image sequence and corresponding scene labels using an artificial intelligence model; wherein the user interface data includes a user operation interface layout and visual elements of the user operation interface; the user operation interface layout includes the spatial arrangement, proportional relationship, and functional division of each component of the user operation interface; and visual elements refer to visual components that do not change over time and constitute a video image; and the device-related information includes operation time information and historical scene recognition information.

[0038] The device comprises:

[0039] a pattern recognition module, configured to obtain user interface data of a current screen image sequence and current device related information, and determine a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship;

[0040] a parameter determination module, configured to determine a corresponding mode weight and a basic threshold value of the usage time of the electronic device according to the current scene label;

[0041] A state determination module is used to determine the working period, continuous use time and actual use time of the electronic device under the current scene tag based on the current device related information;

[0042] A policy decision module is used to generate a time management policy for the electronic device based on the mode weight, the basic usage time threshold, the working period, the continuous usage time and the actual usage time.

[0043] A computer electronic device includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the steps of the electronic device usage time control method based on deep learning are implemented as described above.

[0044] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for controlling the usage time of an electronic device based on deep learning as described above.

[0045] This application has the following advantages:

[0046] In an embodiment of the present application, relative to the problems in the prior art of "inability to identify screen content in a refined manner, high resource consumption, privacy risks, single management strategy, poor user experience, and easy to induce avoidance behavior", the present application provides a solution for screen content recognition based on deep learning and control of usage time based on multi-dimensional factors, specifically: obtaining user interface data and current device-related information of the current screen image sequence, and determining the current scene label corresponding to the user interface data and the current device-related information based on the correspondence; determining the corresponding mode weight and the basic threshold of the usage time of the electronic device based on the current scene label; determining the working period, continuous usage time and actual usage time of the electronic device under the current scene label based on the current device-related information; generating a time control strategy for the electronic device based on the mode weight, the basic threshold of usage time, the working period, the continuous usage time and the actual usage time. Through multi-level deep learning, it can achieve refined recognition of screen content and distinguish different usage modes within the same application; by deploying artificial intelligence models, the main calculations are completed locally to provide privacy protection; through model optimization and hardware adaptation, computing resource consumption is reduced and battery life is extended; through multi-dimensional factors and usage scenarios, differentiated management and control strategies are generated to improve user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0048] Figure 1 This is a flowchart of a method for controlling the usage time of an electronic device based on deep learning provided by an embodiment of the present application;

[0049] Figure 2 This is a schematic diagram of the network structure of a multi-level deep learning model provided by an embodiment of the present application;

[0050] Figure 3 This is a schematic diagram of the network structure of the spatial feature flow and the temporal feature flow provided in one embodiment of the present application;

[0051] Figure 4 This is a flow chart of a feature fusion machine provided in one embodiment of the present application;

[0052] Figure 5 This is a flowchart of a hierarchical classification strategy provided by an embodiment of the present application;

[0053] Figure 6 This is a flow chart of a context awareness module provided in one embodiment of the present application;

[0054] Figure 7 This is a flowchart of a multi-dimensional strategy decision engine provided by an embodiment of the present application;

[0055] Figure 8 This is a flowchart of a real-time duration statistics algorithm provided by an embodiment of the present application;

[0056] Figure 9 This is a flowchart of multi-factor assessment and early warning triggering provided by an embodiment of the present application;

[0057] Figure 10 This is a flow chart of generating a behavior intervention gradient strategy according to an embodiment of the present application;

[0058] Figure 11 This is a structural block diagram of a device for controlling the usage time of an electronic device based on deep learning provided by an embodiment of the present application;

[0059] Figure 12 It is a structural diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0060] To make the objectives, features, and advantages of this application more readily apparent, the present application is further described below in conjunction with the accompanying drawings and specific embodiments. It is apparent that the embodiments described are only a portion of the embodiments of this application, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments in this application without inventive effort are also within the scope of protection of this application.

[0061] By analyzing existing technologies, the inventors discovered that the application package name is the operating system's unique identifier for the application. Its design purpose is to locate the application as a whole, rather than to parse its internal layered content. Simple rules, such as "limit gaming app usage to one hour per day," use static thresholds and fail to incorporate dynamic variables into the control logic. Existing technologies primarily use application package names or simple rules for time control, which presents the following problems:

[0062] a. Insufficient granularity: Existing technologies use package names as the basis for control, essentially treating an "application" as the smallest, indivisible unit. These technologies can only identify the application layer, lacking the ability to identify interface elements and functional modules within an application, and are unable to distinguish between different content types within the same application. For example, existing technologies cannot distinguish between learning content and gaming content within an educational application.

[0063] b. Lack of contextual awareness: Rules and data dimensions are limited, and intelligent adjustments cannot be made based on contextual information such as user behavior, usage scenarios, and time periods.

[0064] c. Single-minded control strategies: A "one-size-fits-all" restriction approach is often adopted. However, teenagers and adults have different abilities to control themselves when gaming. Traditional solutions use the same time limit, which can overly restrict adults and under-regulate teenagers, resulting in a poor user experience and easily leading to user avoidance behavior.

[0065] d. High resource consumption: Traditional solutions require uploading screen content to the cloud for analysis. Each screen analysis requires real-time data upload, which can lead to high response latency in weak network environments. Screen content may contain privacy-sensitive information, and uploading it to the cloud poses privacy risks.

[0066] Therefore, a method is needed that can accurately identify screen content, achieve high efficiency and low power consumption, and provide personalized control. The purpose is to achieve refined recognition of screen content and distinguish different usage modes within the same application; how to achieve efficient and low-power screen content recognition while ensuring privacy and security; and how to generate intelligent and personalized usage time control strategies based on multi-dimensional factors.

[0067] It should be noted that in any embodiment of the present invention, including pattern recognition and policy decision-making, pattern recognition involves a multi-level deep learning model and an adaptive edge inference framework. By adopting a multi-level deep learning model architecture, refined recognition of screen content is achieved. As the core algorithm layer, the multi-level deep learning model is responsible for multi-dimensional feature extraction and pattern recognition of the screen image sequence of the electronic device. The adaptive edge inference framework is a support platform for model deployment and execution. Its core function is to adapt the trained multi-level deep learning model to edge computing devices (such as smartphones, tablets and other terminals) for efficient operation. The adaptive edge inference framework ensures that the model can still meet the real-time requirements on resource-constrained edge devices by dynamically adjusting the model computing resource allocation and optimizing the inference process, while reducing dependence on cloud computing.

[0068] like Figure 2 As shown in the figure, the multi-level deep learning model adopts a dual-stream network structure, including a spatial feature stream and a temporal feature stream. It combines these two features through an attention-guided adaptive fusion mechanism, and introduces a context-aware module to enhance recognition capabilities. Specifically, the multi-level deep learning model includes spatial and temporal feature streams, a feature fusion mechanism, a hierarchical classification strategy, and a context-aware module. A phased training strategy is adopted for the multi-level deep learning model, with the preceding stage providing pre-trained parameters for the subsequent stages. This avoids the convergence difficulties caused by simultaneously optimizing all parameters in a complex network, thereby improving training efficiency and model performance.

[0069] As an example, the multi-level deep learning model is trained in five stages, as follows:

[0070] Phase 1: Spatial Feature Stream Pretraining. The goal is to enable the spatial feature stream to master the ability to extract visual features from single-frame images, such as UI (user interface) layout, icon shape, and text area. This is accomplished by acquiring large-scale sample image datasets, such as ImageNet, to learn basic visual features. This dataset also includes screenshot sample datasets. These screenshot sample datasets can be obtained by capturing various application interfaces and annotating UI element categories, such as buttons, text boxes, and icons.

[0071] The spatial feature flow is based on a lightweight CNN network called MobileNetV3-Small. A sample image dataset is used to pre-train the MobileNetV3-Small backbone network, giving it image classification capabilities. The model's last classifier layer is replaced with a sample screenshot dataset to train it to recognize screen-specific UI element combinations, such as formula editing boxes in educational applications and virtual buttons in games. The backbone network parameters are frozen, and only the classifier layer is fine-tuned to prevent pre-training knowledge from being forgotten.

[0072] Phase 2: Pre-training the temporal feature stream. This aims to enable the temporal feature stream to extract dynamic features from multiple frames, such as the motion trajectories of interface elements and user operation trajectories. This involves acquiring a video action recognition dataset to learn dynamic behavior patterns and a screen operation video dataset, which includes continuous frames of user actions such as sliding and clicking.

[0073] The temporal feature stream is based on the lightweight ShuffleNetV2 network and adds a temporal processing module. The ShuffleNetV2 backbone network is pre-trained using a video action recognition dataset, enabling it to classify video actions, such as swiping, clicking, and page turning. The temporal processing module uses screen operation video data to train the model to recognize dynamic features of these actions, such as consecutive click frames in gaming and page turning frames in document reading. Contrastive learning enhances the robustness of temporal features. For example, after training, the temporal feature stream can distinguish between the high-frequency changes of rapid switching in short videos and the low-frequency changes of slowly turning pages in reading.

[0074] Stage 3: Feature Fusion Module Training. During training, the parameters of the spatial and temporal feature streams are fixed, and only the attention network of the fusion module is trained. Input is a sequence of screen images labeled with usage patterns, along with the spatial and temporal feature vectors of the corresponding sequences. The attention network learns the weights α for spatial features and β for temporal features through gradient descent, such that the product α × spatial features + β × temporal features has the maximum inter-class distance.

[0075] Phase 4: Secondary classifier training. The secondary classifier consists of a primary coarse classifier and a secondary fine classifier. The secondary fine classifier can further subdivide the primary category into subtypes. The training data is a screen content dataset annotated with multiple levels of labels and the comprehensive feature vector output by the fusion module. After the primary classifier outputs the primary coarse category of the screen content, the secondary fine classifier is activated. The training data is the subdivided scenes under the primary coarse category, and each sample is annotated with a secondary label. The secondary classifier parameters are optimized using the cross-entropy loss function so that it can distinguish the feature differences between different scene modes.

[0076] Phase 5: Contextual module training and full network fine-tuning. Contextual information such as device-related time data, device status, and user operation data is integrated into the model to enhance scene understanding capabilities. By inputting screen content samples with contextual annotations, screen image sequences, and corresponding sample datasets, the GRU network encodes temporal and historical features. The encoded contextual features are concatenated with the fused features and input into a hierarchical classifier. The training model adjusts the classification results based on the context. All module parameters are unfrozen, and full network fine-tuning is performed to optimize the contextual information and visual features in a coordinated manner.

[0077] The multi-level deep learning model trained through the above five stages can analyze the patterns between the user interface data and device-related information of the device screen image sequence and the corresponding scene labels, and find the mapping patterns between the user interface data and device-related information of the device screen image sequence and the corresponding scene labels through the self-learning and adaptive characteristics of the artificial intelligence model, so as to obtain the correspondence between the user interface data and device-related information of the device screen image sequence and the corresponding scene labels.

[0078] As an example, the adaptive edge inference framework includes model quantization and optimization, hardware acceleration adaptation, and cloud-edge collaborative inference. Model quantization and optimization deploys deep learning models locally, eliminating the need to upload image data to the cloud. Hardware acceleration adaptation enables efficient local inference, with feature extraction, fusion, and classification of all screen images performed on the device. The cloud-edge collaborative inference mechanism is only used for model updates; daily management and control do not rely on the cloud, and the original screen image is processed entirely locally.

[0079] Reference Figure 1 , shows a method for controlling the usage time of an electronic device based on deep learning provided by an embodiment of the present application;

[0080] The method comprises:

[0081] S110, obtaining user interface data of a current screen image sequence and current device related information, and determining a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship;

[0082] S120: Determine a corresponding mode weight and a basic threshold of the usage time of the electronic device according to the current scene tag;

[0083] S130, determining the working period, continuous usage time, and actual usage time of the electronic device under the current scene tag based on the current device related information;

[0084] S140: Generate a duration management strategy for the electronic device based on the mode weight, the basic usage duration threshold, the working period, the continuous usage duration, and the actual usage duration.

[0085] In the embodiments of the present application, multi-level deep learning is used to achieve refined recognition of screen content and distinguish different usage modes within the same application; by deploying artificial intelligence models, the main calculations are completed locally to provide privacy protection; through model optimization and hardware adaptation, computing resource consumption is reduced and battery life is extended; through multi-dimensional factors and usage scenarios, differentiated management and control strategies are generated to improve user experience.

[0086] Below, a method for controlling the usage duration of an electronic device based on deep learning in this exemplary embodiment will be further described.

[0087] As described in step S110, device-related information and user interface data and device-related information of the current screen image sequence are obtained, and a current scene label corresponding to the user interface data and the current device-related information is determined based on the corresponding relationship; wherein the current device-related information includes time data and historical pattern recognition data.

[0088] It should be noted that the collection of screen images:

[0089] 1. Sampling frequency: Dynamically adjusted according to the device status, once every 30 seconds under normal conditions;

[0090] 2. Resolution processing: Scale the original screen image to 224×224 pixels to reduce the amount of calculation;

[0091] 3. Privacy protection: All image processing is done locally and the original screen image is not uploaded.

[0092] Equipment related information collection:

[0093] 1. Collect current operation time information, such as current time, day of the week, and whether it is a weekday.

[0094] 2. Collect the current device status, such as battery level, network status, list of running applications, etc.;

[0095] 3. Collect user operation data, such as clicks, slides, and other interactive behavior characteristics between users and the screen.

[0096] In one embodiment of the present invention, the specific process of "determining the current scene label corresponding to the user interface data and the current device related information according to the corresponding relationship" in step S110 can be further explained in combination with the following description.

[0097] Determine the spatial characteristics and temporal characteristics of the user interface data and the current device related information as described in the following steps;

[0098] As an example, a single-frame screen image in the current screen image sequence is processed through the spatial feature flow. The spatial feature flow is based on the lightweight CNN network of MobileNetV3-Small to process the single-frame screen image and extract the user interface data in the single-frame screen image. The user interface data includes the user operation interface layout data and its visual elements and other static features.

[0099] The user interface layout refers to the spatial arrangement, proportional relationships, and functional zoning of the user interface components within the current screen image sequence. Component spatial arrangement includes the fixed positions of elements like buttons, menu bars, and progress bars; hierarchical structure includes the nesting of interface elements, such as pop-ups overlaying the main interface; and functional zoning refers to the division of the user interface into distinct areas.

[0100] Visual elements refer to the visual components of a video that do not change over time. Examples include the color and tone of the user interface, the shape and geometry of interface components, interface textures and patterns, text content and fonts, and interface icons and symbols.

[0101] In a specific implementation, Figure 3 As shown, the processing steps of the spatial feature flow include:

[0102] 1. Input layer: 224×224×3 RGB image;

[0103] 2. Backbone network: includes the initial convolutional layer, multiple inverted residual blocks, SE attention module and global average pooling layer;

[0104] 3. Output: 576-dimensional feature vector, encoding the spatial information of the screen.

[0105] As an example, a temporal feature stream processes multiple consecutive screen image frames in the current screen image sequence. The temporal feature stream is based on the lightweight ShuffleNetV2 network and adds a temporal processing module to process multiple consecutive screen image frames to capture dynamic change features. Dynamic change features include the motion trajectory of interface elements and device-related information. The motion trajectory of interface elements refers to the dynamic data of the user interface data extracted by the spatial feature stream. Device-related information is the interaction data between the user and the screen, such as page switching, clicking on the input box, page turning frequency, and operational interaction data of game content.

[0106] In a specific implementation, Figure 3 As shown, the processing steps of the time series feature flow include:

[0107] 1. Input layer: N × 224 × 224 × 3 image sequence (N = 3 to 5);

[0108] 2. Backbone network: includes spatiotemporal convolutional layers, ShuffleNet units, temporal attention modules, and temporal pooling layers;

[0109] 3. Output: 464-dimensional feature vector, encoding the dynamic information of the screen.

[0110] As described in the following steps, a fusion feature is generated based on the spatial feature and the temporal feature.

[0111] As an example, an attention-guided adaptive fusion mechanism is used to dynamically adjust the weights of spatial and temporal features. The feature dimensions of the spatial and temporal features are aligned, and the dimensions of the spatial and temporal features are unified through a fully connected layer (FC). After dimensionality reduction, the features are normalized to ensure the stability of subsequent weight calculations. After preprocessing, the reduced spatial and temporal features are concatenated and subjected to a nonlinear transformation. The spatial and temporal features are then weighted and fused, and the feature dimensions are unified through a fully connected layer to produce the final fused feature vector.

[0112] In a specific implementation, Figure 4 As shown in Figure 2, the feature fusion mechanism includes:

[0113] 1. Input: spatial feature vector (576 dimensions) and temporal feature vector (464 dimensions);

[0114] 2. Attention network: Generates spatial feature weight α and temporal feature weight β (α + β = 1);

[0115] 3. Weighted fusion: α×spatial features + β×temporal features;

[0116] 4. Output: 512-dimensional fused feature vector.

[0117] As described in the following steps, the current scene label is generated based on the fusion features.

[0118] The corresponding scene labels and their confidence levels are output through a hierarchical classification strategy. Specifically, Figure 6 As shown in Figure 2, a coarse-to-fine hierarchical classification strategy is adopted to improve the scene recognition accuracy. The hierarchical classification strategy includes a first-level coarse classifier and a second-level fine classifier.

[0119] The first-level coarse classifier is used to classify screen content into major categories, such as entertainment, study / work, social, and system / other.

[0120] The secondary subclassifier is used to select the corresponding secondary classifier for subclassification based on the primary classification results. For example,

[0121] Entertainment can be divided into: Games-Action, Games-Strategy, Games-Casual, Videos-Long Content, Videos-Short Content, Music, Reading and Entertainment, etc.

[0122] Secondary subclassifiers for learning / work: online courses, educational applications, document editing, programming development, data analysis, meetings / calls, etc.

[0123] Secondary subclassifiers for social networking: instant messaging, social media browsing, social media creation, forums / communities, etc.

[0124] System / Other secondary subclassifiers: system settings, app store, file management, tool applications, etc.

[0125] Finally, the scene label generator combines the first-level and second-level classification results to generate the final scene label and its confidence.

[0126] In another embodiment of the present invention, the specific process of "determining the current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship" in step S110 can be further explained in combination with the following description.

[0127] Determine the spatial characteristics and temporal characteristics of the user interface data and the current device related information as described in the following steps;

[0128] As described in the following steps, generating fusion features based on the spatial features and the temporal features;

[0129] It should be noted that the extraction of spatial features and temporal features, as well as the generation of fusion features can refer to the above description and will not be repeated here.

[0130] Determining enhanced features of the current device-related information as described in the following steps;

[0131] It should be noted that the current device related information includes operation time information and historical scene identification information, and may also include device environment information if necessary. This example uses a context awareness module, such as Figure 6 As shown, the model introduces operation time information, historical scene recognition information, and device environment information context to enhance the model's understanding of usage scenarios. The context-aware module includes a temporal context encoder and a historical context memory module. When device environment information is included, it also includes an environmental context awareness module.

[0132] The temporal context encoder processes time information such as the hour, day of the week, and whether it is a weekday to generate temporal features. The historical context memory module maintains the most recent K pattern recognition results, capturing usage continuity and generating historical features. The environmental context perception module processes environmental information such as screen light intensity, device motion, and device location to generate environmental features. The encoded context features are fused to output enhanced features.

[0133] As described in the following steps, the current scene label is generated according to the fused features and the enhanced features.

[0134] It should be noted that the enhanced features after processing will interact with the fused features. The spatial feature stream mainly extracts spatial features such as UI layout and visual elements of a single-frame image; the temporal feature stream focuses on the dynamic change characteristics of multiple-frame images. Device-related information can provide a richer context for the interpretation of these visual features. For example, when judging a scene containing a screen image of a person's portrait and chat history, the user's interactive behavior information can be combined. If it is detected that the user frequently clicks on the input box and sends messages, the possibility of a social scene is greatly increased.

[0135] As described in step S120, the corresponding mode weight and the basic threshold of the usage time of the electronic device are determined according to the current scene label.

[0136] It should be noted that after the scene labels are identified through the multi-level deep learning model, an intelligent and personalized usage time management strategy is generated based on the multi-dimensional strategy decision engine. Figure 7 As shown, the decision engine includes a parameter definition module, a real-time duration statistics algorithm, a multi-factor evaluation and warning triggering algorithm, a behavior intervention gradient strategy generator, and a personalized and adaptive mechanism. The screen content pattern information of the scene tag is input into the parameter definition module. The parameter definition module, combined with the real-time duration statistics algorithm and the multi-factor evaluation and warning triggering algorithm, calculates the information and feeds it to the behavior intervention gradient strategy generator. The data generated by the intervention gradient strategy generator and processed by the multi-factor evaluation and warning triggering algorithm are then transmitted to the personalized and adaptive mechanism to generate a personalized intervention strategy.

[0137] Parameter definition module, used to define basic parameters and advanced parameters;

[0138] Real-time duration statistics algorithm, such as Figure 8 As shown, it is used to track the duration of various usage patterns;

[0139] Multi-factor assessment and early warning triggering algorithms, such as Figure 9 As shown, it is used to dynamically adjust the threshold and trigger an early warning;

[0140] Behavioral intervention gradient strategy generator, such as Figure 10 As shown, it is used to generate a gradient intervention strategy;

[0141] Personalization and adaptive mechanisms are used to learn from user behavior and optimize decision-making strategies.

[0142] Among them, the basic parameters of the parameter definition module include mode definition and current status;

[0143] The mode definition includes the ID, name, mode weight, and basic threshold of usage time of each usage scenario mode.

[0144] The current status includes the current scene mode, cumulative duration (actual usage time), continuous usage time, and session start time.

[0145] As an example, mode weight is a weighting coefficient assigned to different content types, which is used to adjust the actual usage time. Entertainment content (such as games) has a higher weight, and learning content has a lower weight. Mode weights can be determined according to different content types, or different mode weights can be assigned according to different content types in different time periods. Working hours are time periods divided according to the impact of the length of time electronic devices are used on users within different time ranges, and can be divided into weekday daytime, weekday evening, weekday late night, and weekends. According to Beijing time, weekday daytime can be Monday to Friday 05:00-18:00; weekday evening can be Monday to Friday 18:00-00:00; weekday late night can be Monday to Friday 00:00-05:00.

[0146] For example, when the working hours are 18:00 on weekdays, the mode weight of the game-leisure mode in this working hour is 1.2, the mode weight of the video-short content mode in this working hour is 1.3, the mode weight of the social media-browsing mode in this working hour is 1, and the mode weight of the study / work-education application mode in this working hour is 0.8.

[0147] For example, when the working hours are 02:00 a.m. on weekdays, the mode weight of the game-leisure mode in this working hour is 1.3, the mode weight of the video-short content mode in this working hour is 1.4, the mode weight of the social media-browsing mode in this working hour is 1.1, and the mode weight of the learning / work-education application mode in this working hour is 0.9.

[0148] As an example, the basic usage time threshold is a pre-set initial usage time limit for different content types. For the game-casual mode, the basic usage time threshold is set at 90 minutes; for the high-addiction video-short content mode, the basic usage time threshold is set at 75 minutes.

[0149] The usage time threshold can also be set differently for the same content type based on different work hours. For example, if the work hours are weekdays from 6:00 PM to 12:00 AM, the usage time threshold for Game-Casual mode is set to 90 minutes; if the work hours are weekends, the usage time threshold for Game-Casual mode can be set to 120 minutes.

[0150] For advanced parameters of parameter definition module:

[0151] Policy configuration: including mode switching cooldown time, confidence threshold, etc.;

[0152] Time period coefficient: adjustment coefficient for different time periods;

[0153] Fatigue model: a fatigue coefficient calculation model based on continuous use time;

[0154] User profile: including age group, habit factors, sensitivity to intervention, etc.;

[0155] Response History: A record of the user's responses to past interventions.

[0156] When the scenario tag is input into the parameter definition module of the multi-dimensional strategy decision engine, the parameter definition module determines the corresponding mode weight and usage time basic threshold according to the ID of the usage scenario mode of the current scenario tag.

[0157] As described in step S130, the working period, continuous use time and actual use time of the electronic device under the current scene tag are determined based on the current device related information.

[0158] It should be noted that the current device-related information includes current operation time information, such as time, day of the week, whether it is a weekday, etc. The working period, the user's continuous use time of the device and the actual use time can be calculated based on the current operation time information.

[0159] The working hours are used to determine the time period coefficient, and can also be used to assign different weights to different types of content based on the working hours, or to set different usage time basic thresholds for different types of content based on the working hours.

[0160] When calculating continuous usage time, to avoid statistical distortion of continuous usage time caused by frequent mode switching, a mode switching cooldown period is configured in the advanced parameter mode of the parameter definition module. During the cooldown period, the continuous usage time continues to accumulate during mode switching. If the cooldown period is exceeded, the continuous usage time will be reset to the latest continuous usage time of the current mode.

[0161] Actual usage time refers to the actual usage time of various types of content when users operate electronic devices. It is obtained by recording and counting the start and end time of users' use of the device, and can reflect the users' actual usage of different content types.

[0162] As an example, when calculating the continuous usage time, the time is allocated proportionally within the cooling-down period. When the system detects that the user switches from mode A to mode B, a cooling-down period of fixed length is started. During this period, the time is not immediately attributed to the new mode B, nor is it completely reserved for the old mode A. Instead, this period of time is split proportionally between A and B according to preset rules. Specifically, the user is in mode A and suddenly switches to mode B. The cooling-down period is set to 10 seconds. If the switching time is within 10 seconds, 70% of the time is counted in mode A and 30% in mode B. If the switching time is greater than 10 seconds, the time is fully counted in mode B to confirm that the mode switch is complete.

[0163] As described in step S140, the working period, continuous use time and actual use time of the electronic device under the current scene tag are determined based on the current device related information.

[0164] Reference Figure 8 As shown, in one embodiment of the present invention, the specific process of "determining the working period, continuous usage time and actual usage time of the electronic device under the current scene label based on the current device related information" in step S140 can be further explained in combination with the following description.

[0165] S141. Generate a corresponding dynamic usage duration threshold according to the basic usage duration threshold, the working period, and the continuous usage duration.

[0166] In one embodiment of the present invention, the specific process of step S141 "generating a corresponding dynamic usage duration threshold according to the basic usage duration threshold, the working period, and the continuous usage duration" can be further explained in combination with the following description.

[0167] As described in the following steps, a time period coefficient corresponding to the current scene label is determined according to the working time period.

[0168] It should be noted that the advanced parameters in the parameter definition module include time period coefficients corresponding to working hours. Different time period coefficients are assigned based on the impact of electronic device usage during different working hours, such as weekday daytime, weekday evening, and weekends. Within a 24-hour period, the later the time of day, the lower the time period coefficient.

[0169] As an example, the working period is 7:00-12:00 on weekdays, and the period coefficient is 1.0; the working period is 00:00-2:00 on weekdays, and the period coefficient is 0.8; the working period is 6:00-18:00 on weekends, and the period coefficient is 1.2; the working period is 7:00-11:00 on weekends, and the period coefficient is 1.1.

[0170] As described in the following steps, a fatigue coefficient is determined based on the continuous use time.

[0171] It should be noted that the fatigue coefficient is a quantitative indicator of the user's visual fatigue caused by continuous use of electronic devices, and it increases with the duration of continuous use. The fatigue coefficient is used to dynamically adjust the dynamic threshold of usage time, making the management strategy more consistent with human physiology. For example, the longer the continuous use time, the higher the fatigue coefficient, the lower the dynamic threshold, and the earlier the forced rest period is triggered. The fatigue coefficient is calculated using a fatigue model that supports both exponential and linear models.

[0172] As an example, for the game-casual mode, the user's continuous usage time is 25 minutes, and the fatigue coefficient is 1.1; for the video-short content mode, the user's continuous usage time is 45 minutes, and the fatigue coefficient is 1.3.

[0173] As described in the following steps, the dynamic usage duration threshold is generated according to the basic usage duration threshold, the time period coefficient, and the fatigue coefficient.

[0174] By using the time period coefficient and fatigue coefficient as factors to limit usage time, the larger the product of the time period coefficient and fatigue coefficient, the smaller the dynamic usage time threshold, and the shorter the usage time. The basic usage time threshold is used as the baseline value, and is adjusted by the limiting factor through division to obtain the final dynamic usage time threshold. The calculation formula for the dynamic usage time threshold is as follows:

[0175] Dynamic threshold of usage time = basic threshold of usage time ÷ (time period coefficient × fatigue coefficient)

[0176] By using the basic duration threshold and the calculation of the period coefficient and fatigue coefficient, the fixed duration limit is converted into an adaptive dynamic threshold, so as to carry out personalized management and control, making the management and control strategy more in line with actual usage scenarios and health needs.

[0177] In a specific implementation, the scenario is game-leisure mode, the working period is weekday evening, the period coefficient is 1.0, the basic threshold of usage time is 90 minutes, the continuous usage time is 40 minutes, and the fatigue coefficient is 1.3. Then for this scenario, the dynamic threshold of usage time = 90 ÷ (1.0 × 1.3) ≈ 69 minutes.

[0178] S142: Generate a weighted usage duration according to the mode weight and the actual usage duration.

[0179] By weighting the actual usage time based on the mode weights of different content types, we can obtain the weighted usage time, which can scientifically quantify the usage impact of various types of content. The formula is as follows:

[0180] Weighted usage time = actual usage time × mode weight

[0181] In a specific implementation, the scenario is game-leisure mode, the working period is weekday evening, the mode weight is 1.2, the actual usage time is 50 minutes, and the weighted usage time = 50×1.2 = 60 minutes.

[0182] S143: Generate a duration management policy for the electronic device according to the duration dynamic threshold and the weighted usage duration.

[0183] As an example, the remaining usage time is generated based on the dynamic usage threshold and the weighted usage time; the warning level is determined based on the remaining usage time; and the time management strategy is determined based on the warning level. Specifically, the remaining usage time is determined based on the difference between the dynamic usage threshold and the weighted usage time. The warning levels are divided into four levels, including normal warning level, standard warning level, emergency warning level, and restriction level. The normal warning level means that the remaining usage time is sufficient; the standard warning level means that the remaining usage time is close to the warning time; the emergency warning means that the remaining usage time is very little; and the restriction level means that the dynamic usage threshold has been exceeded.

[0184] For example, if the remaining usage time is ≥ 30% of the dynamic usage time threshold, the system is considered at the normal warning level. For example, if the dynamic usage time threshold is 100 minutes and the weighted usage time is ≤ 70 minutes, and the remaining usage time is ≥ 30 minutes, the system is considered at the normal warning level.

[0185] When 10% of the dynamic usage threshold is less than or equal to the remaining usage time and less than 30% of the dynamic usage threshold, the system enters the standard warning level. For example, if the dynamic usage threshold is 100 minutes, the weighted usage time is between 70 and 90 minutes, and the remaining usage time is between 10 and 30 minutes, the system enters the standard warning level.

[0186] If the remaining usage time is less than 10% of the dynamic usage time threshold, the system enters the emergency warning level. For example, if the dynamic usage time threshold is 100 minutes and the weighted usage time is greater than 90 minutes, and the remaining usage time is less than 10 minutes, the system enters the emergency warning level.

[0187] If the remaining usage time is ≤ 0, that is, the weighted usage time is ≥ the dynamic usage time threshold, the app is considered restricted. For example, if the dynamic usage time threshold is 100 minutes and the weighted usage time is ≥ 100 minutes, the app is considered restricted.

[0188] The time management strategy includes five levels of intervention:

[0189] Level 0: No perception reminder, through the status bar icon or slight animation, the user can choose whether to pay attention to it, without affecting normal use;

[0190] Level 1: Floating prompt: a semi-floating window pops up on the edge of the screen to display the remaining time. It will automatically disappear after a short display. It does not force user interaction and only provides a gentle prompt.

[0191] Level 2: Strong reminder, a full-screen or half-screen reminder dialog box pops up, and the user needs to click the confirmation button to forcibly interrupt the user's current operation process;

[0192] Level 3: Functional restrictions. Instead of directly forcing the user to exit, some functions are disabled, such as disabling in-app purchases in game mode and disabling full-screen in video mode. Only basic functional access is retained, minimizing user experience interruptions while ensuring control.

[0193] Level 4: Forced rest. After the system countdown, it will automatically exit the application, lock the screen, and other operations to force the user to take a break and avoid damage to health caused by excessive use.

[0194] In a specific implementation, at the normal warning level, an imperceptible reminder with an intervention level of 0 is adopted; at the standard warning level, a floating prompt with an intervention level of 1 is adopted; at the emergency warning level, a strong reminder with an intervention level of 2 is adopted; at the restriction level, a function restriction with an intervention level of 3 or a forced rest with an intervention level of 4 is adopted. For example, when the system times out for the first time, it does not directly force a quit, but adopts function restriction. If the user continues to use the system after the function restriction, the system will perform a forced rest.

[0195] As an example, you can set up smart interruption timing to select the best time to intervene and reduce user interruption, such as at the end of a game round or during a video advertisement.

[0196] like Figure 9 As shown, in one embodiment of the present invention, the specific process of step S143 "generating a duration management strategy for the electronic device based on the duration dynamic threshold and the weighted usage duration" can be further explained in combination with the following description.

[0197] It should be noted that to avoid user resistance caused by a "one-size-fits-all" intervention, this embodiment quantifies the user's acceptance of the control strategy based on their personalized characteristics. When generating the duration control strategy, the user's sensitivity to intervention is extracted from the user's profile, and a milder intervention method is adopted for users with high sensitivity.

[0198] As described in the following steps, a user profile is obtained, and the intervention sensitivity of the corresponding user is determined based on the user profile;

[0199] It should be noted that user profiles include factors such as age, habits, and sensitivity to intervention. User profiles are derived from both active user input and passively collected data. Active user input can include age and health status (such as myopia level). Passively collected data includes past user feedback on interventions, historical user behavior (such as commonly used application types, average daily usage time, and frequency of use), and device usage environments, such as the time periods during which the user uses the device.

[0200] As an example, based on the characteristics of the user profile, an intervention sensitivity coefficient is assigned to the user, with the higher the value, the more sensitive it is. The sensitivity coefficient is between 0.5 and 1.5. For example:

[0201] For adolescent users, due to their weaker self-control, the sensitivity coefficient is set at 1.2-1.5, triggering stricter intervention at the same warning level;

[0202] For adult users, the sensitivity coefficient is set to 0.8-1.0, focusing on efficiency management rather than mandatory restrictions;

[0203] For myopic users, the sensitivity factor is set to 1.3.

[0204] As described in the following steps, the remaining usage time is generated according to the dynamic usage time threshold and the weighted usage time.

[0205] Specifically, the remaining usage time is determined according to the difference between the usage time dynamic threshold and the weighted usage time.

[0206] As described in the following steps, determining the warning level according to the remaining usage time;

[0207] It should be noted that the steps for determining the warning level based on the remaining usage time can be referred to the above description and will not be repeated here.

[0208] As described in the following steps, an intervention strategy corresponding to the user is generated according to the intervention sensitivity and the warning level.

[0209] As an example, the intervention sensitivity includes low sensitivity (sensitivity coefficient ≤ 1.0), medium sensitivity (1.0 < sensitivity coefficient < 2.0), and high sensitivity (sensitivity coefficient ≥ 2.0).

[0210] In a specific implementation, when the user has low or medium sensitivity, at the normal warning level, an imperceptible reminder with an intervention level of 0 is used; at the standard warning level, a floating prompt with an intervention level of 1 is used; at the emergency warning level, a strong reminder with an intervention level of 2 is used; at the restriction level, a function restriction with an intervention level of 3 or a forced rest with an intervention level of 4 is used. For example, when the system times out for the first time, it does not force a direct exit, but uses function restriction. If the user continues to use the system after the function restriction, the system will perform a forced rest.

[0211] In one specific implementation, when the user is highly sensitive, no operation is performed at the normal warning level; at the standard warning level, an imperceptible reminder with an intervention level of 0 is used; at the emergency warning level, a floating reminder with an intervention level of 1 is used; at the restricted level, a strong reminder with an intervention level of 2 is used.

[0212] By using user profiles and sensitivity coefficients, we can adapt strategies to different groups of people, avoid "one-size-fits-all" restrictions, and improve user acceptance. Furthermore, we combine this with a dynamic duration threshold response mechanism to achieve precise early warning.

[0213] In one embodiment of the present invention, it further includes:

[0214] Obtaining user feedback on the time management strategy;

[0215] The duration control strategy is optimized based on the feedback information.

[0216] As an example, feedback information provided by users on the intervention strategy is collected, including user response delay time and operation type, etc. The feedback information can be used as data for updating user sensitivity.

[0217] Here are some examples:

[0218] Example 1

[0219] 1. Initial state:

[0220] The user opens an educational app that contains learning content and game content; the system captures a sequence of screen images.

[0221] 2. Pattern Recognition:

[0222] The spatial feature flow analyzes UI layout and element distribution; the temporal feature flow analyzes screen change patterns; the fused features are processed by the classifier; the recognition result: entertainment-game-leisure, confidence level 0.92.

[0223] 3. Strategic Decision-Making:

[0224] Query the game-casual mode parameters: weight 1.2, threshold 90 minutes;

[0225] The current time is weekday evening, and the time period coefficient is 1.0;

[0226] After 25 minutes of continuous use, the fatigue coefficient is 1.1;

[0227] Dynamic threshold: 90 / (1.0×1.1) ≈ 82 minutes;

[0228] Current weighted cumulative duration: 30 × 1.2 = 36 minutes;

[0229] Remaining time: 82-36 = 46 minutes;

[0230] Warning level: Normal.

[0231] 4. Interactive feedback:

[0232] The status bar shows the game icon and the remaining time; continue to monitor usage.

[0233] Example 2

[0234] 1. Initial state:

[0235] The user is watching a short video app; has been using it for 45 minutes.

[0236] 2. Pattern Recognition:

[0237] Recognition result: Entertainment - Video - Short Content, confidence level 0.95.

[0238] 3. Strategic Decision-Making:

[0239] Query video-short content mode parameters: weight 1.3, threshold 75 minutes;

[0240] The current time is late at night, and the time period coefficient is 0.8;

[0241] After 45 minutes of continuous use, the fatigue coefficient is 1.3;

[0242] Dynamic threshold: 75 / (0.8×1.3) ≈ 72 minutes;

[0243] Current weighted cumulative duration: 45×1.3 = 58.5 minutes;

[0244] Remaining time: 72-58.5 = 13.5 minutes;

[0245] Warning level: Standard warning.

[0246] 4. Interactive feedback:

[0247] A floating prompt is displayed: "You have watched the short video for 45 minutes, it is recommended to take a break"; the user clicks "Remind me later"; the system records the response and adjusts the timing of the next reminder.

[0248] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0249] Reference Figure 11 , shows a device for controlling the usage time of an electronic device based on deep learning provided by an embodiment of the present application, wherein the device is used to establish a correspondence between user interface data of a device screen image sequence and device-related information and corresponding scene labels through an artificial intelligence model; wherein the user interface data includes a user operation interface layout and visual elements of the user operation interface, wherein the user operation interface layout includes the spatial arrangement, proportional relationship and functional division of each component of the user operation interface, and the visual element refers to the visual component that does not change over time that constitutes the video screen; the device-related information includes operation time information and historical scene recognition information;

[0250] Specifically include:

[0251] The pattern recognition module 1110 is configured to obtain user interface data of a current screen image sequence and current device related information, and determine a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship;

[0252] A parameter determination module 1120 is configured to determine a corresponding mode weight and a basic threshold value of the usage time of the electronic device according to the current scene label;

[0253] A state determination module 1130 is configured to determine the working period, continuous use time, and actual use time of the electronic device under the current scene tag based on the current device related information;

[0254] The policy decision module 1140 is configured to generate a duration management policy for the electronic device based on the mode weight, the basic usage duration threshold, the working period, the continuous usage duration, and the actual usage duration.

[0255] Regarding the specific definition of an electronic device usage time control device based on deep learning, please refer to the definition of an electronic device usage time control method based on deep learning above, which will not be repeated here. The various modules in the above-mentioned script-based electronic device usage time control device based on deep learning can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0256] In one embodiment, a computer device is provided. The computer device can be a client or a server. The internal structure diagram thereof can be as follows: Figure 12 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The readable storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the readable storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a script-based video editing method.

[0257] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for controlling the usage time of an electronic device based on deep learning in the above embodiment is implemented.

[0258] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a method for controlling the usage time of an electronic device based on deep learning in the above embodiment is implemented.

[0259] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described embodiments. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM).

[0260] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0261] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for controlling the usage time of electronic devices based on deep learning, characterized in that: The method is used to establish a correspondence between user interface data and device-related information of a device screen image sequence and corresponding scene labels through an artificial intelligence model; wherein the user interface data includes a user operation interface layout and visual elements of the user operation interface, the user operation interface layout includes the spatial arrangement, proportional relationship and functional division of each component of the user operation interface, and the visual elements refer to the visual components that do not change over time that constitute the video screen; the device-related information includes operation time information and historical scene identification information; The method comprises: Acquire user interface data of a current screen image sequence and current device related information, and determine a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship; Determining a corresponding mode weight and a basic threshold of the usage time of the electronic device according to the current scene label; Determine the working period, continuous usage time and actual usage time of the electronic device under the current scene tag based on the current device related information; A duration management strategy for the electronic device is generated based on the mode weight, the basic usage duration threshold, the working period, the continuous usage duration and the actual usage duration; specifically, a corresponding dynamic usage duration threshold is generated based on the basic usage duration threshold, the working period and the continuous usage duration; a weighted usage duration is generated based on the mode weight and the actual usage duration; and a duration management strategy for the electronic device is generated based on the dynamic duration threshold and the weighted usage duration.

2. The method according to claim 1, characterized in that The step of determining the current scene label corresponding to the user interface data and the current device related information according to the corresponding relationship includes: Determine spatial characteristics and temporal characteristics of the user interface data and the current device related information; generating a fusion feature based on the spatial feature and the temporal feature; The current scene label is generated according to the fusion feature.

3. The method according to claim 1, characterized in that The step of determining the current scene label corresponding to the user interface data and the current device related information according to the corresponding relationship includes: Determine spatial characteristics and temporal characteristics of the user interface data and the current device related information; generating a fusion feature based on the spatial feature and the temporal feature; determining enhanced features of the current device-related information; The current scene label is generated according to the fused feature and the enhanced feature.

4. The method according to claim 1, wherein The step of generating a corresponding dynamic usage duration threshold according to the basic usage duration threshold, the working period, and the continuous usage duration includes: Determining a time period coefficient corresponding to the current scene label according to the working time period; determining a fatigue factor based on the continuous use duration; The dynamic usage duration threshold is generated according to the basic usage duration threshold, the time period coefficient and the fatigue coefficient.

5. The method according to claim 1, wherein The step of generating a duration management policy for the electronic device based on the duration dynamic threshold and the weighted usage duration includes: Obtaining a user profile, and determining the intervention sensitivity of the corresponding user based on the user profile; generating a remaining usage time according to the dynamic usage threshold and the weighted usage time; Determine the warning level based on the remaining usage time; An intervention strategy corresponding to the user is generated according to the intervention sensitivity and the warning level.

6. The method according to claim 1, characterized in that Also includes: Obtaining user feedback on the time management strategy; The duration control strategy is optimized based on the feedback information.

7. A device for controlling the usage time of an electronic device based on deep learning, characterized in that: The device is used to establish a correspondence between user interface data and device-related information of a device screen image sequence and corresponding scene labels through an artificial intelligence model; wherein the user interface data includes a user operation interface layout and visual elements of the user operation interface, the user operation interface layout includes the spatial arrangement, proportional relationship and functional division of each component of the user operation interface, and the visual elements refer to the visual components that do not change over time that constitute the video screen; the device-related information includes operation time information and historical scene identification information; The device comprises: a pattern recognition module, configured to obtain user interface data of a current screen image sequence and current device related information, and determine a current scene label corresponding to the user interface data and the current device related information based on the corresponding relationship; a parameter determination module, configured to determine a corresponding mode weight and a basic threshold value of the usage time of the electronic device according to the current scene label; A state determination module is used to determine the working period, continuous use time and actual use time of the electronic device under the current scene tag based on the current device related information; A policy decision module is used to generate a time management policy for the electronic device based on the mode weight, the basic usage time threshold, the working period, the continuous usage time and the actual usage time; specifically, to generate a corresponding dynamic usage time threshold based on the basic usage time threshold, the working period and the continuous usage time; to generate a weighted usage time based on the mode weight and the actual usage time; and to generate a time management policy for the electronic device based on the dynamic usage time threshold and the weighted usage time.

8. A computer electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by the processor.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method and device for measuring average power consumption

    CN116338434A

  • Method and system for balancing user experience and power consumption, equipment and storage medium

    CN116954348A