A mobile terminal video backup management method
By implementing differentiated management based on naming number and correlation calculation in mobile terminal video backup, the problems of high resource consumption and low retrieval efficiency in existing technologies are solved, achieving efficient and secure video backup and management.
Patent Information
- Application Number
- CN202511563648.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-10-30
AI Technical Summary
Existing mobile terminal video backup methods generally adopt uniform backup parameters and strategies, resulting in high resource consumption, high storage costs, and low retrieval efficiency when users search for specific content in massive amounts of videos.
Based on the naming and numbering of videos in mobile terminals, valid data is obtained and the correlation is calculated. The data is then classified into different levels, and tag recognition and classification backup are performed separately. The backup strategy is optimized to achieve differentiated management.
By implementing differentiated management, we can reduce computing resource consumption and data transmission volume, lower storage costs, improve retrieval efficiency, ensure timely backup and security protection of important videos, and enhance the overall performance of backup management.
Smart Images

Figure CN121029499B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a mobile terminal video backup management method. BACKGROUND
[0002] With the wide popularity of intelligent mobile terminals, people are used to storing massive video data in devices such as mobile phones and tablets, from daily life records to important work materials, mobile terminal videos carry rich information. However, problems such as device loss, hardware damage and system crash occur frequently, and once an accident occurs, the unbacked video data will be permanently lost, causing great loss to the user. Traditional local backup methods, such as using U disk and mobile hard disk, not only have the risk of damage and loss of backup media, but also need to manually copy when synchronizing data between mobile phones, computers and other devices, which is cumbersome and prone to data version confusion.
[0003] In recent years, cloud storage technology has provided a new solution for mobile terminal video backup management, and users can upload videos to the cloud for unified management and storage. However, existing video backup methods based on cloud storage generally use unified backup parameters and strategies, regardless of video type, shooting device or use scenario, and use the same backup process and storage standard, which results in large consumption of backup management resources and high storage costs, and low search efficiency when users search for specific content in massive videos. SUMMARY
[0004] The problem solved by the present application is to improve the efficiency of mobile terminal video backup management.
[0005] To solve the above problems, the present application provides a mobile terminal video backup management method, comprising:
[0006] Based on the naming number of each video in the mobile terminal, the effective data in the time period from the previous update node to the current node is obtained according to the current update node, wherein the effective data includes timestamp, verification times and preset update duration, and the preset update duration includes the duration from the previous update node to the current node;
[0007] According to the effective data, the association degree of the naming number of each video and the terminal user is obtained, and each video is classified according to the association degree;
[0008] Each of the videos is labeled to obtain a plurality of classification labels, and each of the classification labels is bound to the corresponding time period in the video;
[0009] Backup data corresponding to each video is selected according to the level of each video, the video and the naming number are backed up according to the corresponding backup data, and the backed up device backup video is classified and arranged according to the classification label.
[0010] Optionally, the association degree between the naming number of each video and the terminal user obtained according to the effective data comprises:
[0011] The first coefficient, the second coefficient and the third coefficient are obtained based on the trained reinforcement learning model;
[0012] The first contribution value is obtained according to the timestamp, the verification number and the first coefficient, the second contribution value is obtained according to the timestamp, the verification number and the second coefficient, and the third contribution value is obtained according to the timestamp, the preset update duration and the third coefficient;
[0013] The association degree is obtained according to the first contribution value, the second contribution value and the third contribution value.
[0014] Optionally, the association degree obtained according to the first contribution value, the second contribution value and the third contribution value is represented by a first formula as:
[0015] ,
[0016] Wherein, Hg represents the association degree, α1 represents the first coefficient, α2 represents the second coefficient, α3 represents the third coefficient, Ts represents the timestamp, Tc represents the verification number, and Js represents the preset update duration.
[0017] Optionally, the label identification is performed on each of the videos respectively to obtain a plurality of classification labels, and each of the classification labels is bound to a corresponding time period in the video.
[0018] The pre-trained classification model is used to classify the image and / or text content in each of the videos to obtain a plurality of classification labels.
[0019] Optionally, the mobile terminal video backup management method comprises:
[0020] When a plurality of mobile terminals log in to the same cloud account, the similarity between the device backup videos backed up to the cloud by each of the mobile terminals is obtained;
[0021] The device backup videos are processed by merging, de-duplication and / or retention according to the similarity to obtain cloud backup videos.
[0022] Optionally, the classification label comprises a face label; and the mobile terminal video backup management method further comprises:
[0023] According to the face label, a backup video with the same face label in a historical cloud backup video is obtained, and a to-be-processed video is obtained;
[0024] Identity relationship of the face label is obtained based on scene information and face information in the to-be-processed video, wherein the identity relationship is a relationship between the face label and a terminal user;
[0025] The device backup video is classified and updated according to the identity relationship.
[0026] Optionally, the identity relationship of the face label is obtained based on the scene information and the face information in the to-be-processed video, comprising:
[0027] The scene information is analyzed to obtain a scene label, the face appearance frequency of the face label is obtained according to the to-be-processed video and the historical cloud backup video, and a face analysis is performed on the face information to obtain a face result;
[0028] The identity relationship is obtained according to the face appearance frequency, the scene label and the face result.
[0029] Optionally, the level division of each video according to the correlation degree comprises:
[0030] The correlation degree is compared with a preset first threshold and a preset second threshold,
[0031] When the correlation degree is greater than or equal to the preset first threshold, the level is a high level;
[0032] When the correlation degree is less than the preset first threshold and greater than the preset second threshold, the level is a medium level;
[0033] When the correlation degree is less than or equal to the preset second threshold, the level is a low level.
[0034] Optionally, the backup data corresponding to each video is selected according to the level of each video, comprising:
[0035] When the level is the high level, a high backup frequency, a high intensity encryption and a high level verification are selected as the backup data;
[0036] When the level is the medium level, a medium backup frequency, a medium intensity encryption and a medium level verification are selected as the backup data;
[0037] When the level is the low level, a low backup frequency, a low intensity encryption and a low level verification are selected as the backup data.
[0038] Optionally, before the step of selecting the corresponding backup data according to the level of each video, the method further comprises:
[0039] acquiring a network state, a power condition and a CPU running state of the mobile terminal;
[0040] judging a backup state of the mobile terminal according to the network state, the power condition and the CPU running state;
[0041] if the backup state is a preparation state, executing the step of selecting the corresponding backup data according to the level of each video;
[0042] if the backup state is a waiting state, returning to execute the step of acquiring the network state, the power condition and the CPU running state of the mobile terminal after a preset time interval, until the backup state is the preparation state.
[0043] The beneficial effects of the mobile terminal video backup management method of the present invention are as follows: Based on the naming number of each video in the mobile terminal, valid data within the time period from the previous update node to the current node is obtained according to the current update node, avoiding invalid processing of historical duplicate data, reducing the consumption of computing resources and data transmission volume of the mobile terminal, improving backup preprocessing efficiency, and using the "naming number" as a unique identifier to associate data, ensuring that the update data of the same video is traceable when multiple devices are synchronized, providing a unified and accurate data source for subsequent correlation calculation, and avoiding correlation evaluation deviation caused by data chaos. Based on valid data, the relevance of each video's naming number to the end-user is determined. This relevance is calculated using features such as timestamps, verification counts, and preset update intervals. Each video is then categorized according to its relevance, enabling differentiated management by prioritizing important videos and processing secondary videos as needed. This avoids using high-frequency, high-redundancy backup strategies for low-value videos, reducing cloud storage costs and terminal bandwidth consumption. The categorization results provide a basis for accurate matching of subsequent backup data, ensuring that strongly related videos (such as important meetings and private calls) receive stricter security protection and more timely backup updates, addressing the shortcomings of traditional backups that treat important data and ordinary data indiscriminately. Each video is tagged, resulting in multiple category tags. Each category tag is then bound to a corresponding time period within the video, allowing users to quickly filter target videos by tag and locate video segments by time period, eliminating the need to manually search for corresponding segments in massive or related videos, significantly improving retrieval efficiency. The tag-time period binding supports refined management of video content, meeting users' needs for efficient access to videos in specific scenarios. Based on the rating of each video, corresponding backup data is selected. Videos and their names are backed up according to this data, and the backed-up videos are categorized and organized according to tags. This creates a clear structure for the cloud backup library, allowing users to directly access videos in their respective categories via tags, thus improving the efficiency of backed-up data management. This invention balances backup security, optimized resource allocation, ease of operation, and high management efficiency, comprehensively enhancing the overall performance of mobile terminal video backup management. Attached Figure Description
[0044] Figure 1 This is a flowchart illustrating the mobile terminal video backup management method according to an embodiment of the present invention. Detailed Implementation
[0045] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0046] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0047] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0048] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0049] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0050] It is understood that any part of this application concerning data acquisition or collection has been authorized by the user.
[0051] like Figure 1 As shown in the figure, an embodiment of the present invention provides a mobile terminal video backup management method, including:
[0052] Step S1: Based on the naming number of each video in the mobile terminal, obtain the valid data within the time period from the previous update node to the current node according to the current update node. The valid data includes timestamp, verification count, and preset update duration. The preset update duration includes the duration from the previous update node to the current node.
[0053] Specifically, the mobile terminal can be a mobile phone, a tablet computer, etc., and each video in the mobile terminal can include a life video recorded by the mobile terminal in daily life, a customer service center video call, a remote conference video, a social software video call, a smart home (a visual doorbell), a security system (with a voice intercom function), etc. It should be noted that the naming number of the video can be a custom number named by the terminal user, or a naming number automatically generated according to time. In order to better evaluate the situation of the naming number of the video corresponding to the contact person, it is necessary to count the video generation effective data in a preset time period, and according to the current update node and the previous update node, the effective data is obtained for the naming number of each video, wherein the interval time of the update node is a preset update time length which is set in advance, for example, if the current update node is December 10, 2024 12:00, and the previous update node is November 10, 2024 12:00, all videos in the time period are traversed to obtain the effective data. The timestamp refers to the total call duration data in the preset time period and the video, which can be parsed from the video naming number as a timestamp, or the video duration is counted to obtain the timestamp, for example, in the call video, the timestamp includes the calling duration and the called duration; the verification times refer to the call times of the naming number of the video in the preset update time, including the called times and the calling times, it should be noted that in the video of non-call nature such as the life video recorded by the mobile terminal in daily life, the verification times is 0; the preset update time refers to the set update time, that is, the time interval between two update nodes, which can be set to 30 days, and the video of the mobile terminal is updated and backed up every 30 days. In the specific implementation, the regular expression can be used to extract the string in the file name as the timestamp, and the system self-checking tool or the custom checking algorithm is called to check and count the verification times.
[0054] In addition, in order to ensure the validity of the data, the obtained data needs to be screened and verified. The timestamp needs to meet the specific format requirements and be within a reasonable time range; the verification times need to be positive integers and consistent with the actual situation of the video; and the preset update time should also be within a reasonable interval to avoid abnormal values. If the data is found to be inconsistent with the requirements, the system will be marked and try to reacquire or repair the data. For the data with incorrect timestamp format, the system can prompt the user to manually correct or try to infer the correct time from other related metadata; for the case that the verification times is negative or abnormally large, the verification operation is re-performed.
[0055] Step S2, obtaining the association degree between the naming number of each video and the terminal user according to the effective data, and classifying each video according to the association degree.
[0056] Specifically, according to the effective data obtained as described above, the association degree corresponding to the video naming number is calculated, which refers to the association between the video naming number and the mobile terminal user. The association degree between each video naming number and the terminal user can be calculated using a trained preset video naming number association evaluation model, wherein the preset video naming number association evaluation model can be trained using a DNN network (Deep Neural Network) based on historical effective data, historical naming numbers and historical association degrees. The association degree can also be calculated using a weighted calculation method, wherein the weight coefficients of each feature can be obtained by querying a preset cloud storage technology-based mobile terminal video backup and management platform, or can be dynamically adjusted according to actual conditions to adapt to different use scenarios and user needs. In actual application, the model can be trained and optimized based on a large amount of historical data to improve the accuracy of the association degree calculation. The video naming number association degree can be compared with a threshold to classify the videos into different levels to obtain the naming number association level, which facilitates subsequent hierarchical backup.
[0057] Step S3, respectively identifying the labels of each video to obtain a plurality of classification labels, and binding each classification label with the corresponding time period in the video.
[0058] Specifically, the labels of each video are identified respectively, and various technical means such as image recognition technology and natural language processing technology are used to achieve this. For the image content in the video, image recognition algorithms such as convolutional neural network (CNN) can be used to identify elements such as characters, scenes and objects, and generate corresponding classification labels such as "scene-conference room" and "object-computer". For the audio content in the video, speech recognition technology can be used to convert it into text, and natural language processing technology can be used to extract keywords as labels. Taking the identification of conference videos as an example, the names of the participants can be marked by recognizing the faces in the video, the conference location can be marked by analyzing the images of the conference scene, and the conference theme-related keywords can be extracted as labels by processing the audio content.
[0059] It should be noted that due to, for example, conference video, call video and the like, there is more content within the preset update time, if the terminal user wants to find part of the content, it is necessary to manually adjust the video playing position, if the specific position is forgotten, it is time-consuming and laborious to query. Therefore, after label recognition is performed on each video, the obtained classification label is associated and bound with the corresponding time period, that is, on the time axis of video playing, the starting time and the ending time corresponding to each label are determined, the corresponding relationship between the timestamp and the video frame can be used to accurately determine the time range of the label, and the mapping table of the label time period is established to store and manage the associated information, for example, if "Zhang San" appears in the 0:00-5:00 time period in the video, the "person-Zhang San" label is associated with the time period, and is recorded as {“person-Zhang San”, [0:00, 5:00]}. In subsequent query and management of the video, the related video segment can be quickly located according to the label, such as quickly extracting the decision-making segment in the conference,
[0060] Step S4, selecting corresponding backup data according to the grade of each video, backing up the video and the naming number according to the corresponding backup data, and classifying and arranging the backed-up device backup video according to the classification label.
[0061] Specifically, the grade of the divided video represents the importance of the video to the terminal user, for example, the divided grade is strong association, which means that the video is relatively important to the terminal user, and then high frequency, high encryption and the like are used for backup. The video and the naming number are backed up according to the corresponding backup data, the video data is uploaded to the cloud storage, in the uploading process, the backed-up device backup video is classified and arranged according to the classification label of the video, different folders or database tables can be created in the cloud, and the video is stored in the corresponding position according to the label category. The videos marked as "conference" category are stored under the "conference video" folder, and each folder is further subdivided and stored according to the specific classification label of the video, such as "employee welfare" and "project decision", which facilitates subsequent query and management. In the backup process, the time, version and the like of the backup can also be recorded, so as to trace and recover the data.
[0062] In the embodiment, based on the naming number of each video in the mobile terminal, the valid data in the time period from the previous update node to the current node is obtained according to the current update node, invalid processing of historical repeated data is avoided, the calculation resource consumption and data transmission amount of the mobile terminal are reduced, the backup preprocessing efficiency is improved, the same video update data can be traced when multiple devices are synchronized by taking the naming number as a unique identifier, a unified and accurate data source is provided for subsequent correlation degree calculation, and the correlation degree evaluation deviation caused by data confusion is avoided. According to the valid data, the correlation degree between the naming number of each video and the terminal user is obtained, the correlation degree between the video and the terminal user is calculated through the timestamp, the verification number and the preset update duration, each video is classified according to the correlation degree, the differentiated management of “important video priority protection and secondary video on-demand processing” is realized, the low-value video is prevented from being processed by the high-frequency and high-redundancy backup strategy, the cloud storage cost and terminal traffic consumption are reduced, the classification result provides a basis for the accurate matching of subsequent backup data, ensures that the strongly correlated videos (such as important meetings and private calls) are more strictly protected and more timely updated, and solves the defect that important data and ordinary data are treated equally in the traditional backup. Each video is labeled to obtain multiple classification labels, and each classification label is bound to the corresponding time period in the video, so that the user can quickly filter the target video through the label and locate the video segment through the time period, without manually searching for the corresponding segment in the massive video or related video, and the retrieval efficiency is significantly improved. The binding of the label and the time period provides support for the fine management of the video content and meets the efficient calling demand of the user for the specific scene video. The corresponding backup data is selected according to the level of each video, the video and the naming number are backed up according to the corresponding backup data, and the backed-up device backup video is classified and arranged according to the classification label, so that the cloud backup library structure is clear, the user can directly access the video of the corresponding classification through the label, and the data management efficiency after backup is improved. The embodiment of the present application considers the security of backup, the optimized allocation of resources, the convenience of operation and the efficiency of management, and comprehensively improves the comprehensive performance of the mobile terminal video backup management.
[0063] Optionally, the correlation degree between the naming number of each video and the terminal user according to the valid data comprises:
[0064] The first coefficient, the second coefficient and the third coefficient are obtained based on the trained reinforcement learning model.
[0065] Specifically, reinforcement learning can adapt to different scenarios (such as the correlation degree feature difference of social videos and security videos) by dynamically adjusting the coefficients, improve the matching degree of the coefficients and the actual correlation rule, and adapt to the change of the terminal use habit (such as when the user's interaction frequency of a certain type of video increases, the model can automatically increase the corresponding coefficient weight), solve the problem of insufficient adaptability of fixed coefficients to dynamic scenarios. When implemented, first, a reinforcement learning model is constructed. The state space (S) includes effective data, that is, the fluctuation coefficient (σTs=standard deviation (Ts) / mean (Ts)) of the timestamp (Ts, seconds), the success rate (RTc=effective verification times / total verification times) of the verification times (Tc, times), the continuity (σJs=|Js(t)-Js(t-1)| / Js(t-1)) of the preset update duration (Js, seconds), and the accuracy rate of the historical correlation degree calculation (such as the matching rate of the correlation degree of the previous period and the actual video correlation level), where t represents the update node; the action space (A) is the adjustment amount (such as ±0.05, ±0.1) of the model output α1, α2, α3, which needs to satisfy the dimensional adaptation: for example, the unit of α1 is "1 / (second•times)" (adapted to the product of the timestamp and the verification times), the unit of α2 is "1 / (second 2 •times)" (adapted to the product of the timestamp square and the verification times), and the unit of α3 is "1 / second 2 " (adapted to the product of the timestamp and the preset update duration), to ensure that the subsequent contribution value operation is dimensionless; the value range is: for example, α1, α2, α3 ∈ [0.1, 0.8], to avoid that the coefficients are too small or too large, causing the correlation degree to be distorted. The reward function (R) includes positive rewards: when the coefficients output by the model make the correlation degree match the actual correlation level of the video (such as the interaction frequency of a strongly correlated video and the contingency of a weakly correlated video), the reward value is positively correlated with the matching rate (the higher the matching rate, the greater the reward); negative punishment: when the correlation degree misjudgment leads to unreasonable backup strategy (such as insufficient backup frequency of a strongly correlated video), the punishment value is positively correlated with the misjudgment rate. The pre-constructed reinforcement learning model is trained using historical video data (including Ts, Tc, Js, and artificially labeled correlation levels), and the model learns the mapping rule of the coefficients and the correlation degree through offline training. After deployment, incremental training is performed based on newly collected video data every 24 hours to dynamically optimize α1, α2, and α3 until the reward value is stable (for example, the fluctuation is ≤5% for 100 consecutive iterations) or the iteration reaches the preset maximum number of iterations.
[0066] According to the timestamp, the verification times, and the first coefficient, a first contribution value is obtained, according to the timestamp, the verification times, and the second coefficient, a second contribution value is obtained, and according to the timestamp, the preset update duration, and the third coefficient, a third contribution value is obtained.
[0067] Specifically, the first contribution value is represented as , which comprehensively reflects the superimposed contribution of the correlation of the video in the time dimension and the reliability of the metadata. The greater the sum of the timestamp Ts and the check number Tc, the more significant the timestamp Ts feature of the video (such as the timestamp concentration) and the more check numbers Tc (the more reliable the metadata), the stronger the positive contribution to the correlation Hg after being weighted by the first coefficient a1, and the higher the video correlation Hg.
[0068] The second contribution value is represented as , which embodies the synergistic inhibition effect of the timestamp and the check number. The greater the product of the timestamp Ts and the check number Tc, the stronger the "synergistic effect" of the timestamp Ts and the check number Tc, and the more significant the reverse inhibition effect on the correlation Hg. Meanwhile, the second coefficient a2 amplifies this inhibition effect, that is, the greater the second coefficient a2, the stronger the inhibition of the product of the timestamp Ts and the check number Tc on the correlation Hg, and the lower the video correlation Hg.
[0069] The third contribution value is represented as , which embodies the proportion of the video generation time relative to the preset statistical time length, reflects the benchmark adaptation degree of the time correlation, and the greater the ratio of the timestamp Ts to the preset update time length Js, the higher the proportion of the video generation time in the statistical time length (such as the video being generated in the preset time period). After being adjusted by the third coefficient a3, the stronger the positive contribution to the correlation Hg, that is, the smaller the third coefficient a3, the more significant the influence of the ratio on the correlation Hg, and the higher the video correlation Hg.
[0070] The correlation degree is obtained according to the first contribution value, the second contribution value, and the third contribution value.
[0071] Additionally, the obtained correlation degree can be normalized, that is, the correlation degree Hg is mapped to the interval [0, 1], which is convenient for comparison with, for example, a preset threshold to divide the correlation level.
[0072] Optionally, the correlation degree obtained according to the first contribution value, the second contribution value, and the third contribution value is represented by a first formula as follows:
[0073] ,
[0074] wherein Hg represents the correlation degree, a1 represents the first coefficient, a2 represents the second coefficient, a3 represents the third coefficient, Ts represents the timestamp, Tc represents the check number, and Js represents the preset update time length.
[0075] Optionally, the label identification of each video to obtain a plurality of classification labels, and the binding of each classification label with the corresponding time period in the video include:
[0076] The pre-trained classification model is used to classify the image and / or text content in each video to obtain a plurality of classification labels.
[0077] Specifically, in the model construction phase, for the image content in the video, the classification model can use a convolutional neural network (CNN) architecture (such as ResNet, YOLO) to identify visual features such as people, scenes, and objects; for the text content (such as subtitles, on-screen text, and image text recognized by OCR) in the video, the classification model can use a recurrent neural network (RNN) or a Transformer architecture to extract keywords, topics, and other text features; the model output layer is set to have a multi-label classification function, which supports outputting multiple classification labels at the same time (such as “scene-conference room” and “topic-project discussion”). In the model training phase, the training data set contains a large number of annotated video clips, covering life records, meetings, customer service calls, security monitoring, etc. Each clip is annotated with corresponding image labels (such as people, objects, and scenes) and text labels (such as keywords and event types). The model can be trained using supervised learning, with the cross-entropy loss function used to optimize the parameters until the classification accuracy of the model on the validation set is ≥90% (such as the accuracy of person recognition and scene classification), ensuring the reliability of label recognition. The final trained classification model is obtained. After obtaining the video, frame extraction is performed on the input video (such as extracting a key frame every 2 seconds), which reduces the data volume while retaining the core visual information. Text extraction can be performed on the extracted frames using OCR technology (such as Tesseract) to recognize the text content (such as video subtitles and on-screen dates / locations) and convert it into text data. The key frames are input into the pre-trained image classification model, and the image-related labels (such as “object-laptop” and “scene-office”) are output. The text extracted by OCR is input into the text classification model, and the text-related labels (such as “keyword-product plan” and “event type-work report”) are output. The labels output by the model are de-duplicated and merged (such as merging labels of the same face appearing in consecutive frames into a single face label), and the initial classification label set of the video is generated. Based on the timestamps of the video frames (such as frames extracted at 10 seconds and 12 seconds), the start time and end time of each label are determined. For labels that appear stably in consecutive frames (such as a person appearing continuously from 00:01:00 to 00:05:00), they are bound to the corresponding time period to form a “label-time period” mapping relationship (such as {“person-zhangsan”:[00:01:00,00:05:00]}), and the mapping relationship is stored in the video metadata, supporting subsequent quick positioning of video clips through labels.
[0078] Compared with manual annotation or simple keyword matching, the label recognition accuracy is improved by more than 40%, and batch processing of massive videos is supported, solving the problems of low efficiency and strong subjectivity in traditional methods. At the same time, the image (such as person, scene, object) and text (such as keyword, theme) features are covered, and the generated labels are more consistent with the multi-dimensional attributes of the video content (such as "meeting video" can be labeled with "scene-conference room", "person-participant" and "theme-quarterly summary"), solving the defect that single-dimensional labels cannot completely describe the video content. After binding the label with the time period, users can directly locate the specific segment in the video through the label (such as searching "scene-conference room" can jump to all time periods where it appears), which greatly improves the retrieval efficiency of massive videos, and meets the technical goal of "optimizing video management and improving retrieval convenience" in the document. At the same time, the model is trained by multiple scene data, which can adapt to the features of different types of videos such as life records, remote meetings and security monitoring (such as security videos focusing on "scene-entrance" and "object-stranger" labels, and meeting videos focusing on "theme-discussion content" labels), solving the problem of insufficient adaptability of general label system to specific scenes.
[0079] Optionally, the mobile terminal video backup management method comprises:
[0080] When multiple mobile terminals log in to the same cloud account, the similarity between the device backup videos backed up to the cloud by each mobile terminal is obtained.
[0081] Specifically, it's common for a single user to use two or more mobile devices simultaneously. Videos from these devices can vary significantly, and synchronizing them is cumbersome and time-consuming. Therefore, connecting multiple mobile devices to the same cloud storage account allows for video uploads and backups. When multiple devices log into the same cloud account, the cloud storage system assesses the security status of each device. The system first verifies the authorization status of each device using its unique identifier (e.g., IMEI or SN code) to ensure only trusted devices participate in the backup process. Security checks are performed on each device, including system integrity (e.g., rooted / jailbroken status), software security (e.g., presence of malware), and network environment (e.g., trusted network). Only videos backed up by devices that pass the security check are allowed to proceed to the similarity analysis process. After security testing, videos stored in the cloud can be preprocessed to convert videos backed up from different terminals (such as MP4 and MOV formats) into cloud standard formats (such as H.265 encoding), unifying the video format and eliminating the impact of format differences on similarity calculation. For each video segment, two types of features are extracted: image features, which are extracted by using a convolutional neural network (CNN) to extract feature vectors (such as edges, textures, and object contours) from each frame image (1 frame every 5 seconds); and metadata features, including video naming number, timestamp (generation time), number of verifications (interaction frequency), etc., to form a structured feature set (such as {name number: "Conference 20241001", timestamp: 1696108800, number of verifications: 5}). Then, the similarity is calculated by combining image features and metadata features. The cosine similarity algorithm can be used to calculate the matching degree of the feature vectors of the key frames of the two videos, with a value range of [0, 1], to obtain the image feature similarity. If the naming and numbering are consistent, the timestamp difference is less than or equal to 30 minutes, and the number of verifications is less than or equal to 2, the metadata matching degree can be counted as 1. Otherwise, it is reduced according to the deviation ratio to obtain the metadata similarity. Then, the comprehensive similarity S can be calculated by weighted formula, and a threshold is set. For example, completely consistent: S≥0.9, partially consistent: 0.3≤S≤0.9, completely inconsistent: S<0.3.
[0082] Based on the similarity, the device backup videos are merged and / or deduplicated and / or retained to obtain cloud backup videos.
[0083] Specifically, the device backup video can be updated according to a preset threshold and the similarity of each video. If the two videos are completely consistent, the information is completely the same, and the merging processing can be performed, and only one video is reserved. If the two videos are partially consistent, the associated information of the two videos with the same video naming number is partially inconsistent, such as different video naming number association levels. The two partially similar videos are removed from the duplicate part, and the different parts are sorted. The sorting rule is set in advance according to the user demand. If the two videos are completely inconsistent, the video is reserved.
[0084] The embodiment can reduce the redundant storage of the completely consistent video, retain the complete event chain, and be more suitable for the user's demand for continuous content than pure deletion. In addition, the logical association of the backup videos of different terminals is realized by unified naming number and metadata merging, and the orderliness of the cloud video management is improved.
[0085] Optionally, the classification label includes a face label; and the mobile terminal video backup management method further includes:
[0086] According to the face label, a backup video with the same face label in the historical cloud backup video is obtained, and a to-be-processed video is obtained.
[0087] Specifically, a face recognition model based on a convolutional neural network (such as FaceNet) can be used to extract features of the face in the device backup video, that is, by video frame extraction, key point detection is performed on the face area in the frame, and a 128-dimensional feature vector is generated as a unique face identifier (i.e., a “face label”). The cloud establishes a “face label-video index library” to record the storage path, timestamp, and association level of each face label corresponding to the historical backup video, and supports fast query. When a new video generates a face label such as “face label A”, the cloud system retrieves all backup videos containing “face label A” through the index library to obtain the to-be-processed video.
[0088] Based on the scene information and the face information in the to-be-processed video, an identity relationship of the face label is obtained, wherein the identity relationship is a relationship between the face label and a user of the terminal.
[0089] Specifically, the frequency of occurrence of each scene label in the video to be processed is counted, and then a face library is constructed based on the face information of the user of the terminal, the face feature vector of the user of the terminal is pre-acquired and stored as a reference label, the cosine similarity between the face label (such as “face label A”) in the video to be processed and the reference label is calculated, the preliminary identity is confirmed according to the similarity, if the similarity is greater than a preset value, it is determined as “high similarity” (may be the user himself or a relative), and other face labels are determined as colleagues, friends or strangers. Then, further judgment is made in combination with the frequency of occurrence of each scene label and the face label determination result, for example, when the similarity of face label A and the reference label is high, and the scene label “living room” has a high frequency of occurrence and “certain tourist attraction” has a general frequency of occurrence, it is determined that the identity relationship is “family member”, when the similarity of face label B is low or has no similarity, and the scene label “office” has a high frequency of occurrence and the scene label “home” has a frequency of 0, it is determined that the identity relationship is “colleague”, when the similarity of face label C is high, and the scene label “outdoor” has a high frequency of occurrence and the scene label “living room” has a general frequency of occurrence, it is determined that the identity relationship is “friend”, and when the similarity of face label D is high, and only the scene label “outdoor” exists, the frequency of occurrence is high and there is no high-frequency co-occurrence of faces, it is determined that the identity relationship is “stranger”.
[0090] Further, the identity relationship can also be further indicated according to the contact information (remark name, remark photo) in combination, for example, the identity relationship is family member, the remark photo in the remark information is the same as the face label, and the remark name is mother, and then the identity relationship is updated to mother-daughter.
[0091] According to the identity relationship, the backup video of the device is classified and updated.
[0092] Specifically, after the backup video of the new device is generated, the identity relationship of the face label is automatically classified into the corresponding directory, when the identity relationship changes (for example, “stranger” is updated to “friend” because it appears in the family scene for many times), the cloud automatically migrates the video to a new directory and synchronously updates the backup strategy (for example, from weak association to general association), and at the same time, the user can manually correct the identity relationship, and the system records the correction rule (for example, “user labels ‘Li Si’ as ‘friend’”) for optimizing the subsequent relationship inference model.
[0093] Optionally, the identity relationship of the face label based on the scene information and the face information in the video to be processed comprises:
[0094] The scene information is analyzed to obtain a scene label, the frequency of occurrence of the face label is obtained according to the video to be processed and the historical cloud backup video, and the face information is analyzed to obtain a face result.
[0095] Specifically, key frame extraction is performed on the video to be processed, and the foreground (face, object) and background (environmental area) are separated by an image segmentation algorithm (such as MaskR-CNN), the pixel features (such as color distribution, texture, and landmark objects such as furniture, office equipment, and public facilities) of the background area are extracted, a pre-trained scene recognition model (such as a CNN model based on the Places365 dataset) can be used to classify the background features, and scene labels are output, including but not limited to:
[0096] Home scene (such as "living room" and "bedroom", characterized by home furnishings and personal items);
[0097] Office scene (such as "office" and "meeting room", characterized by desks, projectors, and team environment);
[0098] Public scene (such as "street" and "mall", characterized by dense crowds and public facilities);
[0099] Special scene (such as "in-car" and "outdoor nature");
[0100] In the video to be processed and other historical cloud backup videos, the same face label (based on feature vector matching) is tracked by a multi-target tracking algorithm (such as DeepSORT), the video segments where it appears are recorded, a preset time period (such as 30 days) is set, the total number of appearances is counted, i.e., the total number of videos where the face label appears in the video to be processed and other historical cloud backup videos; the number of co-appearances is counted, i.e., the number of times the face label and the face of the terminal user (preset reference face label) appear together in the same video; the frequency value is calculated, i.e., face appearance frequency = (total number of appearances + 2 x number of co-appearances) / (total number of videos in the preset time period), with a value range of [0, 1], ≥ 0.6 as "high frequency", 0.3-0.6 as "medium frequency", and ≤ 0.3 as "low frequency".
[0101] Facial feature extraction: a face recognition model (such as FaceNet) can be used to analyze the face label, i.e., to extract a 128-dimensional feature vector of the face label, and then calculate the cosine similarity between the feature vector and the facial features of the terminal user, if the similarity ≥ 0.85 is "high similarity", otherwise it is "low similarity"; an age prediction model (such as an age classifier based on ResNet) can also be used to determine the age interval of the face (such as "20-30 years old" and "50-60 years old"), and compare it with the age of the terminal user (such as "similar age" "one generation older" "one generation younger"); the genetic correlation of facial key features (such as eye shape, nose shape, and face shape) is analyzed, and "existence of relative features" or "no relative features" is output; the above dimensions are integrated into a structured facial result, such as "high similarity + similar age + existence of relative features" and "low similarity + large age difference + no relative features".
[0102] According to the face appearance frequency, the scene label and the face result, the identity relationship is obtained.
[0103] Specifically, corresponding classification rules can be set for identity relationship determination, such as:
[0104] Rule 1 (family relationship):
[0105] Trigger condition: high frequency of face appearance + core scene label is "family scene" + face result is "high similarity + similar age / long generation / late generation + existence of relative features";
[0106] Output identity relationship: "direct relatives" (such as parents and children) or "collateral relatives" (such as brothers and sisters, uncles and nephews).
[0107] Rule 2 (colleague / friend relationship):
[0108] Trigger condition: low frequency of face appearance + core scene label is "office scene" or "public leisure scene" + face result is "medium similarity + similar age + no relative features";
[0109] Output identity relationship: "colleague" (office scene) or "friend" (public leisure scene).
[0110] Rule 3 (stranger relationship):
[0111] Trigger condition: low frequency of face appearance + core scene label is "public scene" + face result is "low similarity + large age difference + no relative features";
[0112] Output identity relationship: "stranger".
[0113] In addition, the inferred result can also be automatically compared with historical identity relationship records such as manually annotated relationships by terminal users, and if consistent, it is confirmed; if inconsistent (such as inferred as "stranger" but historical record as "friend"), manual verification is triggered, the user corrects and updates the rule library, and the rule weight is iterated based on new video data periodically (such as adjusting the weight proportion of scene label and frequency), improving the long-term inference accuracy.
[0114] The embodiment sets rules based on quantified frequency, scene and face features, avoids the unexplainability of black box model, verifies multiple dimensional features, makes the identity relationship highly match the real social network of terminal users, and solves the relationship confusion problem caused by "only name annotation" in traditional classification.
[0115] Optionally, the level division of each video according to the correlation degree comprises:
[0116] comparing the correlation degree with a preset first threshold and a preset second threshold,
[0117] when the correlation degree is greater than or equal to the preset first threshold, the level is a high level;
[0118] when the correlation degree is less than the preset first threshold and greater than the preset second threshold, the level is a medium level;
[0119] when the correlation degree is less than or equal to the preset second threshold, the level is a low level.
[0120] Specifically, the preset first threshold (Th1) and the preset second threshold (Th2) are critical values for quantifying the video correlation degree level, determined based on the correlation degree distribution characteristics of historical video data, and satisfy Th1>Th2. The threshold is set by statistical analysis of the Hg distribution law of strong correlation, medium correlation and weak correlation videos. For example: the preset first threshold Th1=0.8 (corresponding to the minimum correlation degree of strong correlation video); the preset second threshold Th2=0.55 (corresponding to the highest correlation degree of weak correlation video). Wherein, the threshold supports user self-defined adjustment (such as modification through cloud management interface), adapts to different scenes (such as personal users can reduce Th1 to 0.7, and enterprise users can increase to 0.85 to strictly distinguish strong correlation videos).
[0121] when Hg≥Th1 (such as Hg≥0.8), it is determined that the video level is a high level;
[0122] when Th2<Hg<Th1 (such as 0.55<Hg<0.8), it is determined that the video level is a medium level;
[0123] when Hg≤Th2 (such as Hg≤0.55), it is determined that the video level is a low level.
[0124] The determination result is bound with the video naming number and stored in the video metadata as the basis for subsequent backup strategy matching.
[0125] Optionally, the selecting the corresponding backup data according to the level of each video comprises:
[0126] when the level is the high level, selecting high backup frequency, high intensity encryption and high level verification as the backup data;
[0127] when the level is the medium level, selecting medium backup frequency, medium intensity encryption and medium level verification as the backup data;
[0128] when the level is the low level, selecting low backup frequency, low intensity encryption and low level verification as the backup data.
[0129] Specifically, the backup frequency refers to the frequency of the mobile terminal video to be backed up to the cloud, including high frequency backup, medium frequency backup or low frequency backup, and the specific frequency can be set by the user as needed, in the embodiment, the backup frequency corresponding to high frequency backup, medium frequency backup and low frequency backup is 1 day, 3 days and 5 days respectively; the encryption level refers to the encryption level, including high strength encryption, medium strength encryption or low strength encryption, low strength encryption can use simple encryption algorithm and short key length, medium strength encryption can use symmetric encryption algorithm as the main part, including some simple combination of asymmetric encryption algorithm, high strength encryption can use complex asymmetric encryption algorithm and advanced encryption technology, the specific encryption method can be set by the user as needed; the access verification mode also has multiple, including password verification, SMS verification code verification, fingerprint identification verification and face recognition verification, which can be selected or combined according to the needs, for example, password verification is low level verification, SMS verification code verification is medium level verification, fingerprint identification verification and face verification are high level verification.
[0130] Optionally, before the step of selecting the corresponding backup data according to the level of each video, the method further comprises:
[0131] obtaining the network state, the power condition and the CPU running state of the mobile terminal;
[0132] judging the backup state of the mobile terminal according to the network state, the power condition and the CPU running state;
[0133] if the backup state is the preparation state, performing the step of selecting the corresponding backup data according to the level of each video;
[0134] if the backup state is the waiting state, after a preset time interval, returning to perform the step of obtaining the network state, the power condition and the CPU running state of the mobile terminal, until the backup state is the preparation state.
[0135] Specifically, when performing backup, power consumption and traffic usage will be increased, and factors to be considered include mobile terminal running condition, and mobile terminal video backup is not an urgent action, and can be performed when the mobile terminal is in good running condition. The network state of the mobile terminal refers to data traffic that can be connected to the network in the mobile network, and refers to network data traffic connected by Wi-Fi and other connection methods. When it is a Wi-Fi network, the remaining traffic is considered to be infinite. The power condition includes sufficient power and insufficient power, and the sufficient power data is recorded as 1, and the insufficient power data is recorded as 0. The network state includes sufficient traffic data recorded as 1 and insufficient traffic data recorded as 0. The CPU running state includes running well data recorded as 1 and running lag data recorded as 0. The network state, power condition and CPU running state are calculated by binary AND operation to obtain mobile terminal backup preparation condition data, including preparation ready data or waiting condition data, and then the mobile terminal backup preparation condition is obtained. When the backup condition of the mobile terminal is in a waiting state, the mobile terminal video backup operation is temporarily not performed, and after waiting for a preset time interval, the mobile terminal backup preparation condition is judged again, and if it is in a preparation state, the mobile terminal video backup is performed.
[0136] Although the present application is disclosed as above, the protection scope of the present application is not limited to this. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and these changes and modifications will fall within the protection scope of the present application.
Claims
1. A mobile terminal video backup management method, characterized by, The method comprises the following steps: Based on the naming number of each video in the mobile terminal, the effective data in the time period from the previous update node to the current node is obtained according to the current update node, wherein the effective data includes the timestamp of the video, the check frequency and the preset update duration, and the preset update duration includes the duration from the previous update node to the current node; According to the effective data, the association degree of the naming number of each video and the terminal user is obtained, and each video is classified according to the association degree; Each video is labeled to obtain multiple classification labels, and each classification label is bound to the corresponding time period in the video; According to the level of each video, the corresponding backup data is selected, the video and the naming number are backed up according to the corresponding backup data, and the backup device backup video is classified and arranged according to the classification label; The classification label includes a face label; the mobile terminal video backup management method further comprises: According to the face label, the backup video with the same face label in the historical cloud backup video is obtained to obtain a to-be-processed video; Based on the scene information and face information in the to-be-processed video, the identity relationship of the face label is obtained, wherein the identity relationship is the relationship between the face label and the terminal user, which includes: The scene information is analyzed to obtain a scene label, the face appearance frequency of the face label is obtained according to the to-be-processed video and the historical cloud backup video, and the face information is analyzed to obtain a face result; According to the face appearance frequency, the scene label and the face result, the identity relationship is obtained; According to the identity relationship, the device backup video is classified and updated.
2. The mobile terminal video backup management method of claim 1, wherein, The association degree of the naming number of each video and the terminal user is obtained according to the effective data, which comprises: Based on the trained reinforcement learning model, first, second and third coefficients are obtained; According to the timestamp, the check frequency and the first coefficient, a first contribution value is obtained, according to the timestamp, the check frequency and the second coefficient, a second contribution value is obtained, and according to the timestamp, the preset update duration and the third coefficient, a third contribution value is obtained; According to the first, second and third contribution values, an association degree is obtained.
3. The mobile terminal video backup management method of claim 2, wherein, The association degree is represented by a first formula according to the first, second and third contribution values: , Wherein, Hg represents the association degree, alpha1 represents the first coefficient, alpha2 represents the second coefficient, alpha3 represents the third coefficient, Ts represents the timestamp, Tc represents the check frequency, and Js represents the preset update duration.
4. The mobile terminal video backup management method of claim 1, wherein, Each video is labeled to obtain multiple classification labels, and each classification label is bound to the corresponding time period in the video, which comprises: The pre-trained classification model is used to classify the image and / or text content in each video to obtain multiple classification labels.
5. The mobile terminal video backup management method of claim 1, wherein, When multiple mobile terminals log in to the same cloud account, the similarity between the device backup videos backed up to the cloud by each mobile terminal is obtained; According to the similarity, the device backup videos are merged, processed and / or de-duplicated and / or retained to obtain cloud backup videos.
6. The mobile terminal video backup management method of claim 1, wherein, The level division of each video according to the correlation degree includes: The correlation degree is compared with a preset first threshold and a preset second threshold, When the correlation degree is greater than or equal to the preset first threshold, the level is a high level; When the correlation degree is less than the preset first threshold and greater than the preset second threshold, the level is a medium level; When the correlation degree is less than or equal to the preset second threshold, the level is a low level.
7. The mobile terminal video backup management method of claim 6, wherein, The selection of corresponding backup data according to the level of each video includes: When the level is the high level, high backup frequency, high intensity encryption and high-level verification are selected as the backup data; When the level is the medium level, medium backup frequency, medium intensity encryption and medium-level verification are selected as the backup data; When the level is the low level, low backup frequency, low intensity encryption and low-level verification are selected as the backup data.
8. The mobile terminal video backup management method of claim 1, wherein, Before the selection of corresponding backup data according to the level of each video, it further includes: The network state, power condition and CPU running state of the mobile terminal are obtained; The backup state of the mobile terminal is determined according to the network state, power condition and CPU running state; If the backup state is a preparation state, the step of selecting corresponding backup data according to the level of each video is performed; If the backup state is a waiting state, after a preset time interval, the step of obtaining the network state, power condition and CPU running state of the mobile terminal is returned to be performed until the backup state is the preparation state.
Citation Information
Patent Citations
A character network relationship discovery and evolution presentation method based on video image recognition
CN109948447A
Person identity recognition method and system, and device and storage medium
WO2024108606A1