Recording automatic control method and apparatus, computer device, and storage medium
By performing multi-dimensional analysis on the video captured by the camera, the recording rate is automatically controlled, solving the inconvenience of users manually adjusting the recording rate, improving the intelligence of video recording and user experience, and reducing the storage of meaningless segments.
Patent Information
- Application Number
- PCT/CN2024/143653
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-02
- Filing Date
- 2024-12-30
- Publication Date
- 2026-02-05
AI Technical Summary
When recording video on existing electronic devices, users need to manually adjust the recording rate, which is inconvenient and may affect video quality. In addition, many meaningless segments are recorded, resulting in a poor user experience.
By performing multi-dimensional analysis of image frames, image sequences, and audio data in the video captured by the camera, its value is evaluated, and the recording rate is automatically controlled to achieve intelligent switching of video recording. This includes object detection, image segmentation, semantic description, and audio recognition, and the recording rate is determined based on the comprehensive value score.
It enables automatic control of video recording, improves user experience, reduces the storage of meaningless clips, saves storage space, and provides a fast playback function.
Smart Images

Figure CN2024143653_05022026_PF_FP_ABST
Abstract
Description
Method and device for automatically controlling recording, computer device and storage medium
[0001] This application is based on and claims priority to Chinese Patent Application No. 202411057720.1, filed on August 2, 2024, the entire contents of which are incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of shooting technology, and more particularly to a recording automatic control method, device, computer device and storage medium. BACKGROUND
[0003] At present, many electronic devices have video recording functions. Common electronic devices with video recording functions include mobile phones, cameras, tablet computers, notebook computers and smart watches, etc. People often use the video recording function of electronic devices to record life and show themselves. However, when people use electronic devices to shoot and record videos, there are often a large number of meaningless segments recorded. For example, when fireworks need to be recorded, the recording often starts before the fireworks start, so that the user needs to manually switch the playback speed of the recorded video from time to time when he wants to view the interesting content, which is more troublesome. Alternatively, the user needs to manually operate the recording rate switching button of the video shooting during the recording process, which is also more inconvenient for the user, and may also cause the electronic device to shake when shooting, affecting the quality of the recorded video and the user experience. SUMMARY
[0004] The technical problem to be solved by the present application is to provide a recording automatic control method, device, computer device and storage medium to automatically switch the recording rate according to the shooting content, realize automatic control of video recording, and improve the user viewing experience.
[0005] In a first aspect, the embodiments of the present application provide a recording automatic control method, which includes: listening to a video collected by a camera; performing value analysis processing on image frames, image sequences and audio data in the listened video respectively to obtain value scores of the image frames, image sequences and audio data respectively; performing comprehensive calculation on the value scores of the image frames, image sequences and audio data to obtain a video comprehensive value; if the video comprehensive value exceeds a first preset threshold, automatically controlling the camera to start recording or recording at a normal rate; if the video comprehensive value is lower than a second preset threshold, automatically controlling the camera to stop recording or recording at a fast rate.
[0006] In a second aspect, the embodiments of the present application also provide a recording automatic control device, which includes units for executing the above method.
[0007] In a third aspect, an embodiment of the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0008] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, wherein the storage medium stores a computer program, and the computer program comprises program instructions, and the program instructions are executed by a processor to implement the method described above.
[0009] Compared with the prior art, the present application respectively performs value analysis processing on the image frames, image sequences and audio data in the monitored video to respectively obtain value scores of the image frames, image sequences and audio data, and performs comprehensive calculation on the value scores of the image frames, image sequences and audio data to obtain a video comprehensive value, and if the video comprehensive value exceeds a first preset threshold, the camera is automatically controlled to start recording or record at a normal speed, and if the video comprehensive value is lower than a second preset threshold, the camera is automatically controlled to stop recording or record at a fast speed. It can be known that the present application can automatically switch the recording speed according to the shooting content to realize automatic control of video recording, that is, the importance of the current shooting content is identified and evaluated by performing value analysis on the perceived video and audio and the like when shooting the video, in order to avoid storing a large amount of useless materials, the camera is automatically controlled to start recording or record at a normal speed when it is evaluated that the importance of the current monitored content is high, and the camera is automatically controlled to stop recording or record at a fast speed when it is evaluated that the importance of the current monitored content is low, and the video recorded and stored can be played quickly when playing, and the video quality is not affected, the user's viewing experience is improved, and the occupied storage space is very small. BRIEF DESCRIPTION OF DRAWINGS
[0010] FIG. 1 is a flowchart of a recording automatic control method provided by an embodiment of the present application;
[0011] FIG. 2 is a sub-flowchart of the recording automatic control method provided by the first embodiment of the present application;
[0012] FIG. 3 is another sub-flowchart of the recording automatic control method provided by the first embodiment of the present application;
[0013] FIG. 4 is another sub-flowchart of the recording automatic control method provided by the first embodiment of the present application;
[0014] FIG. 5 is another sub-flowchart of the recording automatic control method provided by the first embodiment of the present application;
[0015] FIG. 6 is a flowchart of a recording automatic control method provided by a second embodiment of the present application;
[0016] FIG. 7 is a sub-flow diagram of a recording automatic control method according to a second embodiment of the present application;
[0017] FIG. 8 is a schematic block diagram of a recording automatic control device according to a first embodiment of the present application;
[0018] FIG. 9 is a schematic block diagram of a recording automatic control device according to a second embodiment of the present application; and
[0019] FIG. 10 is a schematic block diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0021] It should be understood that, when used in the specification and the appended claims, the terms “comprise” and “include” indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0022] It should also be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clearly indicated by the context, the singular forms “a”, “an” and “the” are intended to include the plural forms as well.
[0023] It should be further understood that, as used in the specification and the appended claims of the present application, the term “and / or” refers to any combination of one or more of the associated listed items and all possible combinations thereof, and includes these combinations.
[0024] Referring to FIG. 1, FIG. 1 is a flow diagram of a recording automatic control method according to a first embodiment of the present application. The recording automatic control method of the present application can be applied to a gimbal camera, a drone with a camera, a terminal device, and the like. As shown in the figure, the method includes the following steps S110-S140.
[0025] S110, listen to the video collected by the camera.
[0026] The video collected in this step can be the video recorded by the camera during shooting, or the video when the camera is previewing. In this application, the automatic control of recording is performed by evaluating the importance of the content of the monitored video, so as to avoid shooting content that is not interesting to the user or content that is not interesting to the user during playback. The content can be quickly played without manual adjustment by the user.
[0027] In S120, the image frames, image sequences and audio data in the monitored video are respectively subjected to value analysis processing to obtain value scores of the image frames, image sequences and audio data.
[0028] In this application, when shooting a video, automatic analysis is performed on the perceived video and audio, etc. to identify whether the current shooting recording / preview content is meaningful, i.e. to evaluate the value of the current shooting recording / preview content.
[0029] Specifically, in some embodiments, as shown in FIG. 2, the value analysis processing of the image frames in the monitored video in step S120 can specifically include steps S121a-S123a.
[0030] In S121a, the image frames in the monitored video are subjected to multi-dimensional analysis processing.
[0031] In this application, the multi-dimensional analysis processing includes two or more of target detection, image segmentation and image semantic description. That is, two or three visual tasks of target detection, image segmentation (including semantic and instance level segmentation) and image semantic description are used to analyze the image content of the image frames in different dimensions in this step.
[0032] In this embodiment, as shown in FIG. 3, the multi-dimensional analysis processing of the image frames in the monitored video in step S121a can specifically include steps S1211a-S1212a.
[0033] In S1211a, the image frames in the monitored video are respectively subjected to target detection and image segmentation processing, and the image content of the image frames is extracted for semantic description of the image.
[0034] In this step, on the one hand, visual detection technology can be used to detect the target of each image frame in the monitored video to obtain the category of each object. Preferably, visual detection technology such as a neural network-based model (e.g. R-CNN and YOLO model, etc.), template matching, etc. can be used to identify and detect the objects existing in the image frame picture to determine the position of the object through the bounding box and determine the category.
[0035] In an aspect, image segmentation can be performed on the image frames. In this embodiment, the image segmentation can include semantic segmentation or instance segmentation, i.e., the image frames can be subjected to semantic segmentation or instance segmentation, so as to extract the scene, the category of the main objects and the relationship therebetween in the image, etc. Preferably, the semantic segmentation can be processed by using a model such as FCN (Fully Convolutional Networks) and U-Net neural network, and the instance segmentation can be processed by using a model such as Mask R-CNN (Mask Region-based Convolutional Neural Network) deep learning model.
[0036] On the other hand, the image frames can also be subjected to semantic-level text description. In this embodiment, a deep visual language model can be used to extract key information from the image and convert it into a natural language description, i.e., the relationship between objects in the image, the behavior occurring, the relationship between objects and the environment, etc. can be described in words according to the image content.
[0037] S1212a, the results of the class of each object and the semantic description of each image frame obtained are filtered by using the preset interest category set and the preset keyword set respectively, so as to obtain the target of interest and the content of interest respectively, and the results after segmentation processing are filtered by using the preset interest scene and category set.
[0038] In this step, the target detection results, the results obtained after image segmentation processing or the results of image semantic description can be filtered according to the pre-set interest category, scene and / or content (keyword) according to the shooting requirements, so as to filter out the objects and categories and scenes that are not of interest to the user, so as to preliminarily know whether the current image frame is an image of interest to the user.
[0039] Specifically, the preset set of interest categories (e.g. people, vehicles, animals, plants, etc.) can be used to filter the categories of all objects obtained by target detection, and the preset set of interest scenes (e.g. basketball game, birthday, Christmas, etc.) and categories (e.g. people, vehicles, objects, etc.) can be used to filter the categories of scenes and main objects obtained after image segmentation processing, and the preset set of keywords (e.g. sports, shooting, running with legs, cake, Christmas tree, etc.) can be used to filter the results of semantic description. For example, when recording a video of watching fireworks, the preset set of interest categories and scenes can be set as fireworks and the moments of the fireworks rising and blooming in the air respectively, and when recording a basketball game video, the preset set of interest categories can be set as some people and a basketball, the preset set of interest scenes can be set as shooting, and the preset set of keywords can be set as jump shot, etc. The results obtained by the three task processing of target detection, image segmentation and image semantic description are filtered by the preset set of interest categories, scenes and keywords, and then the filtered results are analyzed in quality to quickly and accurately evaluate the value of the current image frame in the shooting content, so as to determine whether the image frame in the currently monitored video is of high quality and expected by the user.
[0040] With continuous reference to FIG. 2, S122a, composition quality analysis is performed according to the results of multi-dimensional analysis processing.
[0041] In this application, the multi-dimensional analysis processing can also evaluate whether the image meets the human evaluation standard from the semantic point of view and the composition relationship between the portrait / objects and the background according to the results of multi-dimensional analysis processing.
[0042] Specifically, in this embodiment, the composition quality analysis according to the results of multi-dimensional analysis processing can be: evaluating whether the current image frame of interest of the user meets the preset image composition standard according to the semantic level of the text description of the image content summary and the object position and category obtained by target detection and / or image segmentation processing, and obtaining a composition quality analysis score. For example, when recording a birthday video, the composition quality of the current image frame can also be analyzed according to the positions of the cake and the birthday person in the video image that meets the user's expectation after filtering, and if the preset image composition standard is that the cake and / or the birthday person is in the middle of the image as the standard composition, the composition quality can be evaluated by detecting whether the cake and / or the birthday person is in the middle of the image, etc. It can be understood that the evaluation can be completed by a trained neural network model, and the model can be trained by supervised training through artificially labeled data or by reinforcement learning method through human feedback.
[0043] S123a, obtaining a value score of the image frame according to the composition quality analysis result.
[0044] In this embodiment, the step can specifically be: performing weighted calculation on the composition quality analysis score, and / or obtaining the composition quality analysis score corresponding to each interested category in the image frame and selecting the maximum value or minimum value thereof, and then combining the weighted calculation result and / or the maximum value or minimum value of the composition quality analysis score to obtain the value score of the image frame.
[0045] In this step, the weighted calculation can be defining a weight for each category in the image frame to express its importance in the composition quality analysis, and obtaining a comprehensive weighted value, so as to meet the user's preference for different categories of images and further improve the user experience. Specifically, the comprehensive weighted value can be obtained by adding the weights or taking the maximum value of the weights; and the value score of the image frame obtained by combining the weighted calculation result and / or the maximum value or minimum value of the composition quality analysis score can be: when only one way is used to process the composition quality analysis score, i.e., the composition quality analysis score is weighted calculated or the maximum value / minimum value of the corresponding composition quality analysis score in the image frame is selected, the comprehensive weighted value or the maximum value / minimum value of the composition quality analysis score is taken as the value score of the image frame; and when multiple calculation methods are used to process the composition quality analysis score, i.e., the composition quality analysis score is weighted calculated and the maximum value / minimum value of the corresponding composition quality analysis score in the image frame is selected, the result obtained by adding the comprehensive weighted value and the maximum value / minimum value of the composition quality analysis score is taken as the value score of the image frame, or the maximum value or minimum value of the comprehensive weighted value and the maximum value / minimum value of the composition quality analysis score is taken as the value score of the image frame.
[0046] Preferably, in this embodiment, as shown in FIG. 4, the value analysis processing of the image sequence in the monitored video in step S120 can specifically include steps S121b-S123b.
[0047] S121b, sampling the monitored video according to a preset frame interval to obtain an image sequence.
[0048] In this step, the video stream is sampled according to the preset frame interval, and after N frames are collected (the value of N can be set according to user demand), an image sequence is formed.
[0049] S122b, performing content analysis on the image sequence by using a neural network model to obtain an image sequence score, and scoring the content quality of a video composed of the image sequence collected in a preset time period to obtain a video quality score.
[0050] In this step, the image sequence composed of N frames is input into a neural network model (for example, a Transformer neural network model) for video content analysis. Specifically, one or more neural network models can be used to analyze the image sequence to obtain different evaluation indicators. For example, multiple Transformer neural network models can be used to analyze the behavior, scene, and / or hot spot detection that occurs in the image sequence, and the analysis results can be filtered using pre-set behavior, scene, and / or hot spot types of interest. Understandably, all or part of the evaluation indicators of the behavior, scene, and / or hot spot detection that occurs can be used for evaluation, so as to obtain the image sequence score.
[0051] In this embodiment, video quality evaluation evaluates the quality of the overall video content composed of image sequences from the perspective of the video. Specifically, a trained model can be used to complete the evaluation. The model can be trained using supervised and unsupervised strategies, or reinforcement learning with human feedback (RLHF). The trained model can evaluate the image quality and / or semantic value of the video, so as to obtain the video quality score and obtain high-quality videos.
[0052] S123b, obtaining the value score of the image sequence according to the image sequence score and the video quality score.
[0053] In this step, the image sequence score and the video quality score are combined to evaluate the overall value of the sequence, so as to obtain the value score of the image sequence. First, the image sequence score can be weighted and / or the maximum or minimum value of the image sequence score of different evaluation indicators in the image sequence can be obtained. The preliminary value score of the image sequence can be obtained by combining the weighted calculation result and / or the maximum or minimum value of the image sequence score. The overall value of the sequence can be evaluated by combining the preliminary value score of the image sequence and the video quality score. Understandably, the calculation method of the weighted calculation and / or the maximum or minimum value and the combination of the weighted calculation result and / or the maximum or minimum value of the image sequence score to obtain the preliminary value score of the image sequence is similar to the calculation method used in step S123a for the value score of the image frame. Here, it will not be described again. When performing comprehensive evaluation, the maximum / minimum value of the product of the preliminary value score of the image sequence and the video quality score can be obtained.
[0054] In some embodiments, as shown in FIG. 5, the value analysis and processing of the audio data in the monitored video in step S120 can include steps S121c-S122c.
[0055] S121c, identifying and analyzing the corresponding audio data in the monitored video, and performing audio quality evaluation on the audio data.
[0056] Specifically, in this step, the voice recognition neural network model is used to analyze and identify the corresponding audio data in the monitored video, to obtain the corresponding voice content and the sound type of interest, and the obtained voice content and sound type are filtered by using the preset voice keywords and the set of sound types of interest, and the audio score (i.e., the identification analysis result) is obtained according to the filtering result. If the coincidence rate of the obtained voice content and sound type with the preset voice keywords and the set of sound types of interest exceeds a preset percentage (for example, 70%), the audio score can be set to a higher score, and if it is lower than the preset percentage, it indicates that the content of interest of the user in the audio data is less, and the audio score can be set to a lower score.
[0057] Specifically, in this embodiment, the audio data corresponding to the video composed of the sampled image sequence can be collected by the microphone or the third-party input device, and the collected audio data is input into the neural network model for voice recognition or analysis for further processing to obtain the corresponding voice content, etc. The obtained voice content and sound type can also be filtered by using the preset voice keywords and the set of sound types of interest, to obtain the audio score. At the same time, the quality of the filtered audio data can also be evaluated, for example, the MOS (Mean Opinion Score) score method is used to evaluate the MOS value of the audio data, to obtain the audio quality evaluation result; or the quality of the audio data can also be directly defaulted as full marks.
[0058] S122c, obtaining the value score of the audio according to the identification analysis result and the audio quality evaluation result.
[0059] In this application, the audio score and the audio quality evaluation result of the audio data are comprehensively evaluated to quickly and accurately evaluate the value of the audio data corresponding to the shooting content, so as to determine whether the audio data in the currently monitored video is of high quality and of interest to the user.
[0060] The value score of the audio in this step can be similar to the score calculation method used in step S123b for the value score of the image sequence, which will not be described here.
[0061] S130, the value score of the image frame, the value score of the image sequence, and the value score of the audio data are comprehensively calculated to obtain the video comprehensive value.
[0062] In the present application, the image sequence of the video stream in a certain period of time and the corresponding audio data are collected, and the value of the image frame, the image sequence and the audio data in the period of time is evaluated respectively. After evaluation, the overall value of the video stream in the period of time, i.e. the video comprehensive value, is obtained by multiple comprehensive calculations combined with the value score of the image frame, the value score of the image sequence and the value score of the corresponding audio. In this step, the value score of the comprehensive calculation can be weighted summation or maximum value, etc. Understandably, the score calculation method used in the value score of the image frame in step S123a is similar, which will not be described here.
[0063] S140, if the video comprehensive value exceeds the first preset threshold, the camera is automatically controlled to start recording or record at a normal rate; if the video comprehensive value is lower than the second preset threshold, the camera is automatically controlled to stop recording or record at a fast rate.
[0064] In the present application, the overall value of the video stream in the period of time (i.e. the video comprehensive value) can be used to perform subsequent behavior actions, for example, the recording of the camera can be automatically controlled according to the video comprehensive value. In the present embodiment, when the video comprehensive value exceeds the first preset threshold, it means that the corresponding video is more interesting or meaningful to the user and is interesting to the user, so the camera is automatically controlled to start recording or record at a normal rate; when it is lower than the second preset threshold, it means that the corresponding video is more boring or meaningless to the user and is not interesting to the user, so the camera is automatically controlled to stop recording or record at a fast rate.
[0065] In this step, the second preset threshold is less than the first preset threshold, and the first preset threshold and the second preset threshold constitute the threshold hysteresis range for controlling the recording adjustment of the camera. Preferably, the second preset threshold and the first preset threshold can be set by using a normalization parameter (a value between 0 and 1), for example, the first preset threshold can be 0.7 and the second preset threshold can be 0.6. In the present embodiment, the video comprehensive value can be normalized, and the score range after processing is between 0 and 1. If the normalized video comprehensive value exceeds 0.7, the camera is automatically controlled to start recording or record at a normal rate, while if it is lower than 0.6, the camera is automatically controlled to stop recording or record at a fast rate, and if it is between 0.6 and 0.7, the current recording state is maintained.
[0066] Further, in the embodiment, the fast recording rate can be achieved by keeping the encoding frame rate of the recorded video frames unchanged, reducing the reporting frame rate of the video frames, and / or reducing the frame skipping ratio of the reported video frames, i.e., the frame interval time of the video is lengthened, so that the frame rate is lower than the normal recording frame rate (the normal recording frame rate is usually 24 fps or 30 fps), for example, 1 fps or 0.1 fps. The recording rate is increased by reducing the reporting frame rate of the video frames and / or the frame skipping ratio of the reported video frames, and the video is still played at the frame rate of 24 fps or 30 fps, so that the fast-forward playing effect is achieved.
[0067] As can be seen from the above, the recording automatic control method can analyze the perceived video and audio, etc. to identify whether the current shooting preview content is meaningful, i.e., to evaluate the value or importance of the current shooting preview content. When the importance of the current listening content is high, the camera is automatically controlled to start recording or record at a normal rate, for example, when the results obtained after the image frames in the listened video are processed by target detection, image segmentation and image semantic description are highly matched with the categories, scenes and content of interest set by the user in advance, and the composition quality analysis score, the value score of the image sequence and the value score of the audio are high, the recording automatic control method can automatically control the camera to start recording or record at a normal rate, so that the user can play the video at a normal rate. When the importance of the current listening content is low, the camera can be automatically controlled to stop recording or record at a fast rate, i.e., for the meaningless segments, the recording is usually performed in a time-lapse manner, i.e., the frame interval time of the recording is lengthened, so that the recorded video can be played at an accelerated fast-forward speed, or the recording is automatically stopped to avoid storing a large amount of useless materials. For the meaningful segments, the recording is automatically started and can be recorded at a normal rate, so that the user's viewing experience is improved and the storage space occupied by the overall video is reduced.
[0068] FIG. 6 is a flowchart of a recording automatic control method according to a second embodiment of the present application. As shown in FIG. 6, the recording automatic control method according to the embodiment includes steps S210-S250.
[0069] S210, listening to the video collected by the camera.
[0070] This step is similar to step S110, which will not be described here.
[0071] S220, performing value analysis processing on the image frames, image sequences and audio data in the listened video to obtain the value scores of the image frames, image sequences and audio data, respectively.
[0072] In this step, the value analysis processing is performed on the image frames, the image sequences and the audio data in the monitored video respectively to obtain the value score of the image frames, the value score of the image sequences and the value score of the audio data respectively.
[0073] Specifically, in some embodiments, as shown in FIG. 7, the value analysis processing on the image frames in the monitored video in the step S220 can specifically include steps S221a-S224a.
[0074] S221a, performing multi-dimensional analysis processing on the image frames in the monitored video.
[0075] This step is similar to the step S121a, and will not be repeated here.
[0076] S222a, performing composition quality analysis according to the result of the multi-dimensional analysis processing.
[0077] This step is similar to the step S122a, and will not be repeated here.
[0078] S223a, performing image quality analysis on the image frames in the monitored video according to the preset basic indicators.
[0079] In this step, the preset basic indicators include one or more of brightness, contrast, noise, definition and color. In this application, the image quality of the image frames can be analyzed according to one or more of brightness, contrast, noise, definition and color to obtain the image quality analysis result, i.e. the image quality parameter score.
[0080] S224a, obtaining the value score of the image frames according to the image quality analysis result and the composition quality analysis result.
[0081] In this embodiment, the image frames are comprehensively evaluated in combination with the composition quality analysis result and the image quality analysis result of the image frames, and a more accurate value score of the image frames can be obtained, i.e. the composition quality analysis score and the image quality parameter score are comprehensively evaluated to obtain the value score of the image frames. In this step, when performing comprehensive evaluation, the image quality analysis result and the composition quality analysis result can be multiplied to obtain the maximum value / minimum value, or a model based on neural network can be used for comprehensive evaluation.
[0082] Understandably, the value analysis processing on the image sequences and the audio data in the monitored video in this step is similar to that in the step S120, and will not be repeated here.
[0083] S230, comprehensively calculating the value score of the image frames, the value score of the image sequences and the value score of the audio data to obtain the comprehensive value of the video.
[0084] This step is similar to step S130, and will not be described here.
[0085] S240, if the video comprehensive value exceeds the first preset threshold, automatically control the camera to start recording or record at a normal rate; if the video comprehensive value is lower than the second preset threshold, automatically control the camera to stop recording or record at a fast rate.
[0086] This step is similar to step S140, and will not be described here.
[0087] S250, setting an index mark on the video recorded by the camera according to a preset rule.
[0088] In this application, after obtaining the overall value of the video stream in the preset time period (i.e. the video comprehensive value), the remaining actions can also be performed, for example, in addition to automatically controlling the recording of the camera according to the video comprehensive value, the video can also be marked according to the video comprehensive value.
[0089] In this embodiment, marking the video according to the video comprehensive value can be: setting an index mark on the video recorded by the camera according to a preset rule; specifically, an index mark can be set in the image frame or image sequence or audio data file of the video, and the index mark includes identification information of the image frame or image sequence or audio data file, such as an identification code. It can be understood that the preset rule can be set by the user before recording, for example, the index mark can be set at the image frame or image sequence corresponding to the video stream with the highest video comprehensive value, or the index mark can be set at some preset time points or image frames, i.e. time point marking or frame marking for the video file. After setting the index mark, in subsequent video / image processing, the user can quickly locate, automatically play, crop, generate a time-lapse video, etc. through the index mark.
[0090] FIG. 8 is a schematic block diagram of a recording automatic control device 300 provided by the first embodiment of the application. As shown in FIG. 8, corresponding to the recording automatic control method of the above first embodiment, the application also provides a recording automatic control device 300. The recording automatic control device 300 includes units for executing the recording automatic control method of the above first embodiment, and the device can be configured in a server. Specifically, referring to FIG. 4, the recording automatic control device 300 includes a listening unit 301, an analysis processing unit 302, a comprehensive calculation unit 303, and a recording control unit 304.
[0091] The listening unit 301 is configured to listen to the video captured by the camera; in this embodiment, the video captured by the camera can be the video in the camera shooting recording, or the video in the camera shooting preview.
[0092] The analysis processing unit 302 is configured to perform value analysis processing on the image frames, image sequences and audio data in the monitored video respectively to obtain value scores of the image frames, image sequences and audio data respectively. Specifically, in the embodiment, the analysis processing unit 302 includes an image frame evaluation unit 3021, an image sequence evaluation unit 3022 and an audio evaluation unit 3023.
[0093] The image frame evaluation unit 3021 is configured to perform multi-dimensional analysis processing on the image frames in the monitored video, perform composition quality analysis according to the result of the multi-dimensional analysis processing, and obtain the value score of the image frames according to the result of the composition quality analysis. The multi-dimensional analysis processing includes two or more of target detection, image segmentation and image semantic description. That is, two or three visual tasks of target detection, image segmentation (including semantic and instance level segmentation) and image semantic description are used to analyze the image content of the image frames in different dimensions. In the embodiment, three visual tasks of target detection, image segmentation (including semantic and instance level segmentation) and image semantic description are used to analyze the image frames. Specifically, the image frame evaluation unit 3021 can perform target detection and image segmentation processing on the image frames in the monitored video respectively, and extract the image content of the image frames to perform semantic description of the image. The obtained class of each object and the result of semantic description of each image frame are filtered by using a preset set of interest class and a preset set of keywords respectively to obtain the target of interest and the content of interest respectively, and the result of the segmentation processing is filtered by using a preset set of interest scene and class to obtain the result of interest of the user. For example, when recording a concert video, the set of interest class and the set of interest scene can be set as person and stage performance respectively. Understandably, if the user expects to record the scene of the target person stretching the other arm backward while holding the other arm up, the preset keywords can be arm up and arm backward. The above-mentioned set of interest class, set of interest scene and keywords are used to filter the shooting content, so that the video finally started to record and recorded at a normal speed is the video expected by the user. After the multi-dimensional analysis processing, the result of the multi-dimensional analysis processing is used to evaluate whether the image meets the human evaluation standard from the semantic angle and the composition relationship between the portrait / object and the background to obtain a composition quality analysis score, and the value score of the image frames is obtained according to the composition quality analysis score.
[0094] The image sequence evaluation unit 3022 is configured to sample the monitored video according to a preset frame interval to obtain an image sequence, and score the image sequence to obtain a value score of the image sequence. Preferably, the image sequence evaluation unit 3022 is configured to sample the monitored video according to a preset frame interval to obtain an image sequence, analyze the image sequence by using a neural network model to obtain an image sequence score, score a content quality of a video composed of image sequences collected in a preset time period to obtain a video quality score, and obtain the value score of the image sequence according to the image sequence score and the video quality score.
[0095] The audio evaluation unit 3023 is configured to analyze the corresponding audio data in the monitored video, evaluate the audio quality of the audio data, and obtain a value score of the audio according to the analysis result and the audio quality evaluation result. The audio evaluation unit 304 of the present application analyzes the corresponding audio data in the monitored video by using a speech recognition neural network model to obtain corresponding speech content and a sound type of interest, filters the obtained speech content and sound type by using a preset speech keyword and a sound type set of interest, obtains an audio score (i.e., the analysis result) according to the filtering result, simultaneously evaluates the quality of the audio data, and obtains the value score of the audio by comprehensively considering the analysis result and the audio quality evaluation result.
[0096] The comprehensive calculation unit 303 is configured to comprehensively calculate the value score of the image frame, the value score of the image sequence, and the value score of the audio data to obtain a video comprehensive value. In the present embodiment, the image sequence and the corresponding audio data of the video stream in a monitored time period are collected, the value of the image frame, the image sequence, and the audio data in the time period is evaluated, and the value score of the image frame, the value score of the image sequence, and the value score of the corresponding audio are combined to obtain the overall value of the video stream in the time period, i.e., the video comprehensive value, by multiple times of comprehensive calculation. Preferably, the value score calculation mode of the comprehensive calculation can be weighted summation or maximum value taking, etc.
[0097] The recording control unit 304 is configured to automatically control the camera to start recording or record at a normal speed when the video comprehensive value exceeds a first preset threshold, and automatically control the camera to stop recording or record at a fast speed when the video comprehensive value is lower than a second preset threshold. In this application, the second preset threshold is less than the first preset threshold, and the first preset threshold and the second preset threshold form a threshold hysteresis range for adjusting the recording of the camera. Preferably, in this embodiment, the video comprehensive value can be normalized, the first preset threshold can be set to 0.7, and the second preset threshold can be set to 0.5. If the normalized video comprehensive value exceeds 0.7, the camera is automatically controlled to start recording or record at a normal speed. If the normalized video comprehensive value is lower than 0.5, the camera is automatically controlled to stop recording or record at a fast speed. If the normalized video comprehensive value is between 0.5 and 0.7, the current recording state is maintained. In this embodiment, recording at a fast speed can be achieved by keeping the encoding frame rate of the recorded video frames unchanged, reducing the reporting frame rate of the video frames, and / or reducing the frame skipping ratio of the reported video frames. That is, the frame interval time of the video is lengthened, so that the frame rate is lower than the normal recording frame rate. By reducing the reporting frame rate of the video frames and / or the frame skipping ratio of the reported video frames, the recording speed is increased, and the video is still played at a normal frame rate during playing, which can achieve a fast-forward playing effect.
[0098] It can be seen that, in this application, the importance of the monitored video content is evaluated to automatically control the recording, so that the user can quickly play the content that is not interesting to the user during playing without manual adjustment, which can improve the user experience.
[0099] FIG. 9 is a schematic block diagram of a recording automatic control device 300 according to a second embodiment of the present application. As shown in FIG. 9, the recording automatic control device 300 according to the second embodiment is different from the first embodiment in that a marking unit 305 is added and the specific function of the image frame evaluation unit 3021 is different, and the rest of the structure is the same or similar. In the second embodiment, the image frame evaluation unit 3021 is configured to perform multi-dimensional analysis and processing on the image frames in the monitored video, perform composition quality analysis according to the results of the multi-dimensional analysis and processing, analyze the image quality of the image frames in the monitored video according to preset basic indicators, and obtain a value score of the image frames according to the image quality analysis result and the composition quality analysis result. The preset basic indicators include one or more of brightness, contrast, noise, definition, and color. In the present application, the image quality of the image frames can be analyzed according to one or more of brightness, contrast, noise, definition, and color to obtain an image quality parameter score. In the second embodiment, the image frames are comprehensively evaluated in combination with the composition quality analysis result and the image quality analysis result of the image frames, and a more accurate value score of the image frames can be obtained. The marking unit 305 is configured to set an index mark for the video recorded by the camera according to a preset rule, so that in subsequent video / image processing, the index mark can be used for quick positioning, automatic playing, clipping, generating a time-lapse video, and the like.
[0100] It should be noted that the specific implementation process of the recording automatic control device 300 and each unit can be clearly understood by those skilled in the art, and can be referred to the corresponding description in the foregoing method embodiments. For the convenience and brevity of description, the details are not described herein.
[0101] The recording automatic control device 300 described above can be implemented in the form of a computer program, which can run on a computer device as shown in FIG. 10.
[0102] Referring to FIG. 10, FIG. 10 is a schematic block diagram of a computer device according to an embodiment of the present application. The computer device 500 can be a server, which can be a standalone server or a server cluster composed of multiple servers.
[0103] Referring to FIG. 10, the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501, wherein the memory can include a non-volatile storage medium 503 and an internal memory 504.
[0104] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which when executed, can cause the processor 502 to perform a recording automatic control method.
[0105] The processor 502 is configured to provide computing and control capabilities to support the operation of the entire computer device 500.
[0106] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503, which, when executed by the processor 502, causes the processor 502 to perform an automatic recording control method.
[0107] The network interface 505 is configured to perform network communication with other devices. Those skilled in the art can understand that the structure shown in FIG. 10 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 500 to which the scheme of the present application is applied. Specifically, the computer device 500 can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0108] The processor 502 is configured to run the computer program 5032 stored in the memory to implement the following steps: listening to a video collected by a camera; performing value analysis processing on image frames, image sequences and audio data in the listened video respectively to obtain value scores of the image frames, the image sequences and the audio data respectively; performing comprehensive calculation on the value scores of the image frames, the value scores of the image sequences and the value scores of the audio data to obtain a video comprehensive value; if the video comprehensive value exceeds a first preset threshold, automatically controlling the camera to start recording or recording at a normal rate; if the video comprehensive value is lower than a second preset threshold, automatically controlling the camera to stop recording or recording at a fast rate.
[0109] In an embodiment, the value analysis processing on the image frames in the listened video specifically includes: performing multi-dimensional analysis processing on the image frames in the listened video; performing composition quality analysis according to the result of the multi-dimensional analysis processing; obtaining the value score of the image frames according to the result of the composition quality analysis; wherein the multi-dimensional analysis processing includes two or more of target detection, image segmentation and image semantic description.
[0110] In an embodiment, the multi-dimensional analysis processing on the image frames in the listened video at least includes: performing target detection on the image frames in the listened video to detect objects existing in the image frames and obtain the category of each object, and extracting the image content of the image frames to perform text description on the images; filtering the category of each object and the text description of each image frame obtained by using a preset interest category set and a preset keyword set respectively to obtain interested targets and interested contents respectively.
[0111] In an embodiment, the multi-dimensional analysis on the image frames in the monitored video further comprises: image segmentation processing on the image frames in the monitored video, and filtering the segmented results by using a preset scene and category set of interest.
[0112] In an embodiment, the processor 502, after implementing the composition quality analysis according to the results of the multi-dimensional analysis, further implements the following steps: image quality analysis on the image frames in the monitored video according to a preset basic index, and obtaining a value score of the image frames according to the image quality analysis result and the composition quality analysis result; wherein the preset basic index comprises one or more of brightness, contrast, noise, definition, and color.
[0113] In an embodiment, the value analysis on the image sequence in the monitored video specifically comprises: sampling the monitored video according to a preset frame interval to obtain the image sequence, performing content analysis on the image sequence by using a neural network model to obtain an image sequence score, scoring the content quality of a video composed of the image sequence collected in a preset time period to obtain a video quality score, and obtaining a value score of the image sequence according to the image sequence score and the video quality score.
[0114] In an embodiment, the value analysis on the audio data in the monitored video specifically comprises: recognition analysis on the corresponding audio data in the monitored video, and audio quality evaluation on the audio data; and obtaining a value score of the audio according to the recognition analysis result and the audio quality evaluation result.
[0115] In an embodiment, the recognition analysis on the corresponding audio data in the monitored video specifically comprises: recognition analysis on the corresponding audio data in the monitored video by using a speech recognition neural network model to obtain corresponding speech content and a sound type of interest, and filtering the obtained speech content and sound type by using a preset speech keyword and a set of sound types of interest.
[0116] In an embodiment, the processor 502 further implements the following steps: setting an index mark on the video recorded by the camera according to a preset rule.
[0117] It should be understood that, in the embodiments of the present application, the processor 502 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0118] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the above-mentioned embodiments.
[0119] Therefore, the present application also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. The program instructions are executed by a processor to make the processor perform the following steps: listening to a video collected by a camera; performing value analysis processing on image frames, image sequences and audio data in the listened video respectively to obtain value scores of the image frames, the image sequences and the audio data respectively; performing comprehensive calculation on the value scores of the image frames, the value scores of the image sequences and the value scores of the audio data to obtain a video comprehensive value; if the video comprehensive value exceeds a first preset threshold, automatically controlling the camera to start recording or recording at a normal rate; if the video comprehensive value is lower than a second preset threshold, automatically controlling the camera to stop recording or recording at a fast rate.
[0120] It can be understood that, in some embodiments, the processor can also perform other steps of the recording automatic control method described in the above-mentioned embodiments when executing the program instructions stored in the storage medium. For details, please refer to the above-mentioned embodiments, which will not be described here. The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, and various computer-readable storage media that can store program codes.
[0121] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0122] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and actual implementation can have another division manner. For example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed.
[0123] The steps in the method embodiments of the present application can be adjusted, combined and deleted in sequence according to actual needs. The units in the apparatus embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0124] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a client, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application.
[0125] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for automatically controlling recording, characterized by, The recording automatic control method comprises: listening to the video collected by the camera; value analysis processing is performed on the image frames, image sequences and audio data in the listened video respectively to obtain value scores of the image frames, image sequences and audio data respectively; comprehensive calculation is performed on the value scores of the image frames, image sequences and audio data to obtain a video comprehensive value; if the video comprehensive value exceeds a first preset threshold, the camera is automatically controlled to start recording or record at a normal speed; if the video comprehensive value is lower than a second preset threshold, the camera is automatically controlled to stop recording or record at a fast speed.
2. The recording automation control method of claim 1, wherein, The value analysis processing on the image frames in the listened video comprises: multi-dimensional analysis processing is performed on the image frames in the listened video; composition quality analysis is performed according to the results of the multi-dimensional analysis processing; the value score of the image frames is obtained according to the composition quality analysis result; The multi-dimensional analysis processing comprises two or more of target detection, image segmentation and image semantic description.
3. The recording automation control method of claim 2, wherein, The multi-dimensional analysis processing on the image frames in the listened video at least comprises: target detection is performed on the image frames in the listened video to detect the objects existing in the image frame and obtain the category of each object, and the image content of the image frame is extracted for semantic description of the image; the obtained category of each object and the semantic description of each image frame are filtered by using a preset interest category set and a preset keyword set respectively to obtain the interested target and the interested content respectively.
4. The recording automation control method of claim 3, wherein, The multi-dimensional analysis processing on the image frames in the listened video further comprises: image segmentation processing is performed on the image frames in the listened video, and the segmented results are filtered by using a preset interest scene and category set; wherein the image segmentation processing comprises semantic segmentation or instance segmentation.
5. The recording automation control method of claim 2, wherein, After the composition quality analysis according to the results of the multi-dimensional analysis processing, the image quality of the image frames in the listened video is further analyzed according to a preset basic index, and the value score of the image frames is obtained according to the image quality analysis result and the composition quality analysis result; wherein the preset basic index comprises one or more of brightness, contrast, noise, definition and color.
6. The recording automation control method of claim 1, wherein, The value analysis processing on the image sequences in the listened video comprises: image sequences are obtained by sampling the listened video according to a preset frame interval; content analysis is performed on the image sequences by using a neural network model to obtain an image sequence score, and content quality of a video composed of the image sequences collected in a preset time period is scored to obtain a video quality score; the value score of the image sequences is obtained according to the image sequence score and the video quality score.
7. The recording automation control method of claim 1, wherein, The value analysis processing on the audio data in the listened video comprises: recognition analysis is performed on the corresponding audio data in the listened video, and audio quality assessment is performed on the audio data; the value score of the audio is obtained according to the recognition analysis result and the audio quality assessment result.
8. The recording automation control method of claim 7 wherein, The identifying and analyzing the corresponding audio data in the monitored video specifically comprises: utilizing a speech recognition neural network model to identify and analyze the corresponding audio data in the monitored video, to obtain corresponding speech content and a sound type of interest, and utilizing a preset speech keyword and a sound type of interest set to filter the obtained speech content and sound type.
9. The recording automation control method of claim 1, wherein, The automatic recording control method further comprises: Indexing the video recorded by the camera according to a preset rule.
10. An automatic recording control device characterized by comprising: Comprise: A monitoring unit configured to monitor a video collected by a camera; An analysis processing unit configured to respectively perform value analysis processing on image frames, image sequences, and audio data in the monitored video, to respectively obtain value scores of the image frames, the image sequences, and the audio data; A comprehensive calculation unit configured to comprehensively calculate the value scores of the image frames, the value scores of the image sequences, and the value scores of the audio data, to obtain a comprehensive value of the video; A recording control unit configured to automatically control the camera to start recording or record at a normal speed if the comprehensive value of the video exceeds a first preset threshold, and automatically control the camera to stop recording or record at a fast speed if the comprehensive value of the video is lower than a second preset threshold.
11. The recording automation control apparatus of claim 10, wherein The analysis processing unit comprises: An image frame evaluation unit configured to perform multi-dimensional analysis processing on the image frames in the monitored video, and perform composition quality analysis according to the results of the multi-dimensional analysis processing, and obtain value scores of the image frames according to the results of the composition quality analysis; wherein the multi-dimensional analysis processing comprises two or more than two of target detection, image segmentation, and image semantic description; An image sequence evaluation unit configured to perform value analysis processing on the image sequences in the monitored video, to obtain value scores of the image sequences; An audio evaluation unit configured to perform value analysis processing on the audio data in the monitored video, to obtain value scores of the audio data.
12. The recording automation control apparatus of claim 11, wherein The image frame evaluation unit is specifically configured to: perform target detection on the image frames in the monitored video, to detect objects existing in the image frames, and obtain categories of each object, and extract image content of the image frames, to perform semantic description on the images; respectively utilize a preset category set of interest and a preset keyword set to filter the obtained categories of each object and the semantic description of each image frame, to respectively obtain targets of interest and contents of interest, to realize multi-dimensional analysis processing on the image frames in the monitored video.
13. The recording automation control apparatus of claim 12, wherein The image frame evaluation unit is further configured to perform image segmentation processing on the image frames in the monitored video, and utilize a preset scene set of interest and a category set to filter the results of the segmentation processing, to realize multi-dimensional analysis processing on the image frames in the monitored video; wherein the image segmentation processing comprises semantic segmentation or instance segmentation.
14. The recording automation control apparatus of claim 11, wherein The image frame evaluation unit is further configured to analyze image quality of the image frames in the monitored video according to preset basic indexes, and obtain a value score of the image frames according to the image quality analysis result and the composition quality analysis result, wherein the preset basic indexes include one or more of brightness, contrast, noise, definition and color.
15. The recording automation control apparatus of claim 11, wherein The image sequence evaluation unit is specifically configured to sample the monitored video according to a preset frame interval to obtain an image sequence, analyze content of the image sequence by using a neural network model to obtain an image sequence score, score content quality of a video composed of the image sequence collected in a preset time period to obtain a video quality score, and obtain a value score of the image sequence according to the image sequence score and the video quality score.
16. The recording automation control apparatus of claim 11, wherein The audio evaluation unit is specifically configured to perform recognition analysis on the corresponding audio data in the monitored video, perform audio quality evaluation on the audio data, and obtain a value score of the audio according to the recognition analysis result and the audio quality evaluation result.
17. The recording automation control apparatus of claim 16, wherein The audio evaluation unit is specifically configured to perform recognition analysis on the corresponding audio data in the monitored video by using a speech recognition neural network model to obtain corresponding speech content and a sound type of interest, and filter the obtained speech content and sound type by using a preset speech keyword and a set of sound types of interest, so as to realize the recognition analysis on the corresponding audio data in the monitored video.
18. The recording automation control apparatus of claim 10, wherein The recording automatic control device further includes a marking unit configured to set an index mark for the video recorded by the camera according to a preset rule.
19. A computer device, comprising: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method in any one of claims 1-9 when executing the computer program.
20. A storage medium, characterized by The storage medium stores a computer program, the computer program includes program instructions, and the program instructions can implement the method in any one of claims 1-9 when executed by a processor.
Citation Information
Patent Citations
Video recording device and method
CN102780869A
T video shooting method and terminal equipment
CN109819171A
Video recording frame rate control method and related device
CN113411528A
Image shooting method and device, computer equipment and storage medium
CN115802153A
Automatic recording control method and device, computer equipment and storage medium
CN118984365A