Video processing method and related apparatus
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MASHANG CONSUMER FINANCE CO LTD
- Filing Date
- 2024-11-22
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the traceability process is not convenient enough and is inefficient, making it difficult to effectively improve users' perception of goods or food.
By acquiring video of the object's activity area, determining the video acquisition strategy and reading the prompt information, using a video encoder to encode the video and information, generating video features and information features, inputting them into the model for biological behavior recognition, and generating traceability videos that meet the traceability conditions.
It enhances the flexibility and accuracy of biometric behavior recognition, generates rich and effective traceability videos, and improves the convenience and efficiency of the traceability process.
Smart Images

Figure CN119697429B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a video processing method and related apparatus. Background Technology
[0002] With the continuous development and progress of the Internet and information technology, more and more services are conducted online. As living standards continue to improve, users have higher and higher requirements for everyday items and food. Based on this, in order to enhance users' perception of the items or food they use, traceability has become an important way for providers or recommenders to make recommendations to users. How to improve the convenience and efficiency of traceability is a key focus for providers, recommenders, and users. Summary of the Invention
[0003] In a first aspect, embodiments of this application provide a video processing method, including:
[0004] Acquire the video of the active area of the object, determine the video acquisition strategy corresponding to the video of the active area, and read the prompt information of the video acquisition strategy;
[0005] The video encoder performs video encoding on the video of the active area to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features;
[0006] The video features and information features are input into the model for biological behavior recognition to obtain the recognition result;
[0007] If the identification result satisfies the tracing conditions corresponding to the video acquisition strategy, a tracing video of the object is generated based on the video of the activity area.
[0008] As can be seen in this embodiment, during video processing, after acquiring the video of the object's active area, the video acquisition strategy corresponding to the active area video is first determined. Then, the prompt information of the video acquisition strategy is read, thereby associating the active area video with the video acquisition strategy. The prompt information of the video acquisition strategy prompts subsequent biometric behavior recognition, improving the flexibility and effectiveness of biometric behavior recognition of the active area video. Based on this, the video encoder performs video encoding on the active area video to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features. Then, the video features and information features are input into the model for biometric behavior recognition to obtain the recognition result. To differentiate the results, before performing biometric behavior recognition through the model, the activity area video and prompt information are encoded separately to obtain accurate encoding results. The encoded video features and information features are then used as input to enable the model to generate behavior recognition, thereby improving the accuracy of the model's biometric behavior recognition. After using the model to perform biometric behavior recognition on the activity area video corresponding to the video acquisition strategy and obtaining the recognition results, the source tracing conditions corresponding to the video acquisition strategy influence the generation of source tracing videos. Thus, biometric behavior recognition and source tracing video generation are performed based on the video acquisition strategy, improving the richness and effectiveness of the generated source tracing videos.
[0009] Secondly, embodiments of this application provide a video processing apparatus, including:
[0010] The video acquisition module is used to acquire video of the active area of the object, determine the video acquisition strategy corresponding to the active area video, and read the prompt information of the video acquisition strategy;
[0011] The encoding module is used for the video encoder to encode the video of the active area to obtain video features, and the information encoder to encode the prompt information to obtain information features.
[0012] The biometric behavior recognition module is used to input the video features and the information features into the model to perform biometric behavior recognition and obtain recognition results.
[0013] If the identification result meets the tracing conditions corresponding to the video acquisition strategy, then the video generation module is run. The video generation module is used to generate a tracing video of the object based on the video of the activity area.
[0014] Thirdly, embodiments of this application provide a video processing apparatus, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the video processing method described in the first aspect.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium for storing computer-executable instructions, which, when executed by a processor, implement the video processing method as described in the first aspect.
[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the video processing method as described in the first aspect. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A schematic diagram illustrating the implementation environment of a video processing method provided in this application embodiment;
[0019] Figure 2 A flowchart of a video processing method provided in this application embodiment;
[0020] Figure 3 This is a schematic diagram of video encoding processing provided in an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of text encoding processing provided in an embodiment of this application;
[0022] Figure 5 A schematic diagram illustrating the encoding and biometric behavior recognition process provided in this application embodiment;
[0023] Figure 6 A flowchart illustrating a video processing method for generating traceability videos of farmed animals, provided in this embodiment of the application;
[0024] Figure 7 A flowchart illustrating another video processing method for generating traceability videos of farmed animals, provided in this application embodiment;
[0025] Figure 8 A schematic diagram of a video processing device provided in an embodiment of this application;
[0026] Figure 9 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application. Detailed Implementation
[0027] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments in this specification, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this application.
[0028] In practical applications, in order to improve the perception of goods or food, the production process of goods or the breeding process of ingredients can be traced to provide safer and healthier goods or food.
[0029] To achieve convenient and effective source tracing, this application provides a video processing method. By performing biometric behavior recognition and source tracing video generation on videos of an object's activity area, the source tracing video of the object can be retrieved during the source tracing process, thereby achieving fast and realistic object source tracing. Specifically, in the process of performing biometric behavior recognition and source tracing video generation on videos of an object's activity area, after obtaining the video of the object's activity area, the method first reads the prompt information corresponding to the video acquisition strategy for the video of the activity area. Then, the video of the activity area is encoded using a video encoder to obtain video features, and the prompt information is encoded using an information encoder to obtain information features. The video features and information features are then input into a model for biometric behavior recognition to obtain recognition results. Finally, if the recognition results meet the source tracing conditions corresponding to the video acquisition strategy, the source tracing video of the object is generated based on the video of the activity area. This method improves the flexibility and richness of processing by performing biometric behavior recognition on videos of the activity area according to the video acquisition strategy, and enhances the accuracy and effectiveness of biometric behavior recognition by using a model.
[0030] The video processing methods provided in one or more embodiments of this specification are applicable to the implementation environment of source tracing video processing of objects, such as... Figure 1 As shown, the implementation environment includes at least server 101, which can be one or more servers, a server cluster consisting of several servers, or a cloud server of a cloud computing platform.
[0031] In addition, the implementation environment may also include a video encoder 102, an information encoder 103 and a model 104, wherein the video encoder 102 is used for video encoding, the information encoder 103 is used for information encoding, and the model 104 is used for biometric behavior recognition.
[0032] The implementation environment may also include a video acquisition component 105, which may be a zoomable, rotatable camera. The video acquisition component may also support the configuration of video acquisition strategies. The video acquisition component 105 may be configured in the active area of an object to acquire video of the active area of the object.
[0033] In this implementation environment, the video acquisition component 105 acquires video of the active area of the object. After the server 101 obtains the video of the active area of the object acquired by the video acquisition component 105, it first determines the video acquisition strategy corresponding to the video of the active area and reads the prompt information of the video acquisition strategy. Then, the video encoder 102 performs video encoding on the video of the active area to obtain video features, and the information encoder 103 performs information encoding on the prompt information to obtain information features. The video features and information features are then input into the model 104 for biometric behavior recognition to obtain recognition results. Finally, if the recognition results meet the traceability conditions corresponding to the video acquisition strategy, the traceability video of the object is generated based on the video of the active area.
[0034] Figure 2 This is a flowchart of a video processing method provided in this embodiment, referring to... Figure 2 The video processing method provided in this embodiment specifically includes steps S202 to S208:
[0035] Step S202: Obtain the video of the active area of the object, determine the video acquisition strategy corresponding to the video of the active area, and read the prompt information of the video acquisition strategy.
[0036] The object in this embodiment can be a produced item or a domesticated animal; the domesticated animal can be a intensively raised organism, such as a free-range chicken. For the domesticated animal, generally, an activity area is set up to allow it to move freely. In this embodiment, the activity area video can refer to video filmed within the domesticated animal's activity area. For the item, generally, a production area is set up for the item's production. In this embodiment, the activity area video can refer to video filmed within the item's production area.
[0037] For any product, whether produced manually or by machine, there are specific production and rest periods. Similarly, domesticated animals have both activity and rest periods. To avoid the continuous video capture by video acquisition components deployed in the activity area, which could lead to invalid video captured during extended rest periods and cause wear and tear on the video acquisition components, this embodiment configures a video acquisition strategy. This strategy allows the video acquisition components to capture video according to a specific strategy, thereby improving the usability of the captured activity area video and saving on video acquisition components and video storage space. This embodiment allows for the configuration of video acquisition strategies to meet different traceability requirements.
[0038] Taking free-range chickens as an example, in configuring video acquisition strategies, a strategy for capturing videos of chickens leaving the coop can be configured: continuous video capture from 7:00 AM to 10:00 AM, with checks every hour on the hour; or a strategy for capturing videos of chickens returning to the coop can be configured: continuous video capture from 4:00 PM to 7:00 PM, with checks every hour on the hour; or a close-up video acquisition strategy can be configured: 5 minutes of close-up video capture every hour from 7:30 AM to 7:30 PM. In addition to configuring the capture time and duration, the video acquisition strategy can also configure prompts. For example, the prompt for the chicken leaving the coop video acquisition strategy can be: "Is the chicken leaving the coop behavior present?"; the prompt for the chicken returning to the coop video acquisition strategy can be: "Is the chicken returning to the coop behavior present?"; or, for both strategies, the prompt can be: "Are there chickens currently present?"; the prompt for the close-up video acquisition strategy can be: "Are there chickens currently present? If so, summarize interesting behaviors of the chickens leaving the coop." It should be noted that the above-described video acquisition strategies for chickens leaving the coop, returning to the coop, and close-up video acquisition are merely illustrative. In practical applications, video acquisition strategies may include any one or more of the above three strategies, and may also include other configured video acquisition strategies. This embodiment does not impose any limitations on these strategies. Optionally, the video acquisition strategy may include video acquisition time, video acquisition duration, prompt information, and / or traceability conditions. It should also be noted that the configuration of video acquisition strategies for items is similar to the above and can be configured according to actual needs. This embodiment will not elaborate further on this.
[0039] In practice, after obtaining the video of the active area of the object captured by the video capture component, the video capture strategy corresponding to the active area video is first determined based on the capture time and / or capture duration of the active area video, and then the prompt information of the video capture strategy is read.
[0040] In step S204, the video encoder performs video encoding on the video of the active area to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features.
[0041] In this embodiment, after obtaining the video of the activity area and the prompt information, biometric behavior recognition is performed based on the video of the activity area and the prompt information. In order to improve the accuracy and effectiveness of the large language model in biometric behavior recognition, in this step, the video of the activity area is first encoded to obtain video features, and the prompt information is encoded to obtain information features.
[0042] In practical implementation, due to the differences between video and information, to ensure the effectiveness of encoding, this embodiment provides an optional implementation whereby the video encoder performs video encoding on the active area video to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features. That is, the active area video is input into the video encoder for video encoding to obtain video features, and the prompt information is input into the information encoder for information encoding to obtain information features. The following details the process of the video encoder performing video encoding on the active area video to obtain video features and the process of the information encoder performing information encoding on the prompt information to obtain information features.
[0043] (1) The video encoder performs video encoding on the video of the active area to obtain video features;
[0044] In specific implementation, during the process of video encoder encoding the video of the active area to obtain video features, the video of the active area needs to be input into the video encoder for video encoding to obtain video features. In other words, the active area video is input into the video encoder for video encoding. In order to improve the encoding efficiency of video encoding while ensuring the validity of the video, in an optional implementation method provided in this embodiment, the following operation is performed during the process of inputting the active area video into the video encoder for video encoding to obtain video features:
[0045] Video frames are extracted from the video of the active area to obtain a video frame sequence;
[0046] The video frame sequence is input into the video encoder for video encoding to obtain the video features.
[0047] Specifically, the video of the active area can be skipped to obtain a video frame sequence, which is then input into a video encoder for video encoding to obtain video features.
[0048] In the specific execution process of video encoding by a video encoder, in order to pay more attention to the continuous information of the video, each video frame in the video frame sequence can be encoded as an independent video feature. However, since the size of a video frame is h*w*c, where h, w, and c represent the height, width, and number of channels of the video frame, respectively, the size of the video frame is huge. Directly inputting the video frame into the embedding layer for video frame transformation would involve a large amount of computation. Therefore, in the process of video encoding the video frame sequence, the video encoder can first perform dimensionality reduction processing on each video frame to obtain a dimensionality-reduced video frame sequence, and then perform video encoding based on the dimensionality-reduced video frame sequence.
[0049] In one optional implementation method provided in this embodiment, the video encoder can perform video encoding in the following manner:
[0050] Position encoding is performed on the video frame sequence to obtain position features, and dimensionality reduction processing is performed on each video frame in the video frame sequence to obtain multiple dimensionality-reduced video frames. Then, video frame encoding is performed on each dimensionality-reduced video frame to obtain initial video frame features.
[0051] The positional features of each video frame and the features of the initial video frame are fused to obtain the fused features of each video frame;
[0052] The fusion features of each video frame are input into an autoregressive model for autoregressive calculation to obtain the video frame code of each video frame as the video code.
[0053] Specifically, positional encoding is introduced to enable one-to-one matching of video frames and prompts.
[0054] For example, such as Figure 3 As shown, the video encoder includes a dimensionality reduction model (dimensionality reduction module), a video frame coding model (embedding layer), a position coding model (position coding module), a feature fusion model (feature fusion module), and an autoregressive model (autoregressive module); wherein, the input video frame sequence can be Where n is the number of video frames in the video frame sequence; then, on the one hand, the dimensionality reduction model can use the SVD (Singular Value Decomposition) dimensionality reduction method to reduce each video frame to a t*c (t<(h*w)) matrix, after dimensionality reduction processing Dimensionally reduced video frames The input video frame encoding model, i.e., the embedding layer, performs video frame encoding operations and outputs initial video features. Matrix;
[0055] On the other hand, the positional coding model performs positional coding on a video frame sequence. In the positional coding process, the following is first defined:
[0056]
[0057]
[0058] Where PE is the position encoding matrix initialized with all zeros, which is represented as the maximum position matrix, and P0 is the normalized set of 1 to n;
[0059] To enable the PE to have location encoding information, the following calculation is performed:
[0060]
[0061]
[0062]
[0063]
[0064] Here, div represents the variation coefficient, norm is the normalization function, and then the odd-numbered columns PE(:, 0::2) and even-numbered columns PE(:, 1::2) of the PE matrix are assigned corresponding inverse trigonometric functions to form the corresponding encoding tensors. The dimension of PE becomes PE ;
[0065] Obtaining initial video features After combining the location features (PE), the feature fusion model superimposes the location features and the initial video features to obtain the fused feature X1. ; where X1 ;
[0066] After obtaining X1, input X1 into the autoregressive model for autoregression calculation. The autoregression calculation is as follows:
[0067]
[0068]
[0069]
[0070] Further calculation of the fraction matrix:
[0071]
[0072] Reuse
[0073]
[0074] Calculate and obtain video features It should be noted that the calculations obtained... .
[0075] (2) The information encoder encodes the prompt information to obtain information features;
[0076] The prompt information can be prompt text, and correspondingly, the information encoder can also be a text encoder.
[0077] In the specific execution process, positional encoding is also introduced during information encoding to improve the effectiveness of information encoding. In an optional implementation method provided in this embodiment, the information encoder can encode the prompt information in the following way:
[0078] The prompt information is position-encoded to obtain position features, and the prompt information is initial information-encoded to obtain initial information features;
[0079] The initial information features and the location features are fused, and the fused features are input into an autoregressive model for autoregressive calculation to obtain the information features.
[0080] For example, such as Figure 4 As shown, the text encoder includes a text encoding model (embedding layer), a positional encoding model (positional encoding module), a feature fusion model (feature fusion module), and an autoregressive model (autoregressive module). First, it obtains the prompt text. Where N is 1 and n is the number of texts in the prompt text, the prompt text X is first input into the text encoding model, i.e., the embedding layer, for text conversion to obtain the initial text features. , where 512 is the embedding dimension;
[0081] To imbue text vectors with richer positional information, positional information is introduced:
[0082]
[0083]
[0084] Where PE is the position encoding matrix initialized with all zeros, which is represented as the maximum position matrix, and P0 is the normalized set of 1 to n;
[0085] To enable the PE to have location encoding information, the following calculation is performed:
[0086]
[0087]
[0088]
[0089]
[0090] Here, div represents the variation coefficient, norm is the normalization function, and then the odd-numbered columns PE(:, 0::2) and even-numbered columns PE(:, 1::2) of the PE matrix are assigned corresponding inverse trigonometric functions to form the corresponding encoding tensors. The dimension of PE becomes PE ;
[0091] In obtaining initial text features After combining the location features (PE), the feature fusion model superimposes the location features and the initial video features to obtain X1= ; where X1 ;
[0092] After obtaining X1, input X1 into the autoregressive model for autoregression calculation. The autoregression calculation is as follows:
[0093]
[0094]
[0095]
[0096] Further calculation of the fraction matrix:
[0097]
[0098] Reuse
[0099]
[0100] Calculate and obtain text features It should be noted that the calculations obtained... .
[0101] The above describes in detail the video encoding process of the video encoder and the information encoding process of the information encoder. To improve the encoding accuracy of the trained video encoder and information encoder, in this embodiment, the video encoder and information encoder can be trained together. During training, at the initialization stage, a definition is defined... = As the model is continuously trained, As the model parameters are continuously updated, finally through The loss is calculated using the loss function; specifically, in one optional implementation provided in this embodiment, the video encoder and information encoder are trained in the following manner:
[0102] Read video samples and corresponding prompt information samples from the active area;
[0103] The video frame sequence corresponding to the video sample of the activity area is input into the video encoder to be trained for video encoding to obtain the sample video features, and the prompt information sample is input into the information encoder to be trained for information encoding to obtain the sample information features.
[0104] The training loss is calculated based on the sample video features and sample information features, and the parameters of the video encoder and the information encoder to be trained are updated based on the training loss.
[0105] Specifically, the video encoder and information encoder to be trained are trained in the manner described above until the encoder converges or reaches the preset number of training iterations, thereby obtaining the video encoder and information encoder.
[0106] Step S206: Input the video features and the information features into the model for biological behavior recognition to obtain the recognition result.
[0107] In the steps described above, the video encoder encodes the video of the active area to obtain video features, and the information encoder encodes the prompt information to obtain information features. Based on this, in this step, the video features and information features are input into the model for biometric behavior recognition to obtain the recognition result. The model can be a large language model, such as an LLM (Large Language Model). For example, as... Figure 5 As shown, the prompt text is input into a text encoder for text encoding to obtain text features. A video frame sequence consisting of at least one video frame is input into a video encoder for video encoding to obtain video features. The text features and video features are then input into a model (e.g., LLM) for biometric behavior recognition to obtain the recognition result. Placeholders can also be input into the model to define the structure and shape of the output.
[0108] In practical implementation, the recognition results of biometric behavior identification vary depending on the prompt information. That is, the corresponding recognition results also differ depending on the video acquisition strategy. Taking a chicken leaving the coop video acquisition strategy as an example, if the prompt information is: "Does the chicken leave the coop behavior exist?", the corresponding recognition result is either yes or no. Taking a close-up video acquisition strategy as an example, if the prompt information is: "Does a chicken currently exist? If yes, summarize the chicken's interesting behaviors?", the corresponding recognition result is yes, the chicken's interesting behavior is chasing, or no chicken exists. In other words, the large language model uses different methods to perform biometric behavior identification under different prompt information, obtains the recognition results, and outputs them. In one optional implementation method provided in this embodiment, biometric behavior identification can be implemented in the following way:
[0109] Perform semantic recognition based on the information features to obtain semantic recognition results;
[0110] Based on the semantic recognition results, a recognition result matching the video features is generated.
[0111] Specifically, during the process of biological behavior recognition, the model performs semantic recognition based on information features. After obtaining the semantic recognition result, it can generate a recognition result matching the video features according to the processing strategy corresponding to the semantic recognition result. For example, if the semantic recognition result is "chickens leaving the coop," it calculates whether free-range chickens are included based on video features. If yes, the recognition result is "yes"; otherwise, it is "no." Or, if the semantic recognition result is "chickens returning to the coop," it calculates whether free-range chickens are included based on video features. If yes, the recognition result is "no"; otherwise, it is "yes." Furthermore, if the semantic recognition result is "interesting behavior," it calculates whether free-range chickens are included based on video features. If no, the recognition result is "no"; if yes, it calculates the interesting behavior based on video features and outputs interesting behaviors such as jumping and / or running. In addition to including interesting behaviors, the recognition result can also include the time of occurrence of the interesting behavior; correspondingly, the video description information of the generated source video can also include the time of occurrence and the interesting behavior.
[0112] Step S208: If the recognition result satisfies the tracing conditions corresponding to the video acquisition strategy, generate a tracing video of the object based on the video of the activity area.
[0113] The traceability videos include videos recording the behavior of the animals or videos recording the production progress of the items.
[0114] In practice, video features and information features are input into the model for biological behavior recognition. After obtaining the recognition results, it is necessary to check whether the recognition results meet the traceability conditions corresponding to the video acquisition strategy. Since different video acquisition strategies have different traceability conditions and different detection methods, the process of generating traceability videos based on activity area videos also differs. Continuing with the example of the chicken leaving the coop video acquisition strategy, the close-up video acquisition strategy, and the chicken returning to the coop video acquisition strategy; the essence of the chicken leaving the coop video acquisition strategy and the chicken returning to the coop video acquisition strategy is to obtain relevant videos of chickens leaving and returning to the coop. A large language model can be used to identify whether chickens are leaving or returning to the coop in the acquired activity area videos. The video of chickens returning to the coop is relevant, so the video acquisition strategies for chickens leaving the coop and returning to the coop can be considered as the same type of video acquisition strategy for determining the source conditions. The essence of the close-up video acquisition strategy is to obtain interesting behaviors of free-range chickens. Interesting behaviors contained in the activity area video can be identified by using a large language model. Therefore, the close-up video acquisition strategy can be considered as a separate type for determining the source conditions. Taking the video acquisition strategies for chickens leaving the coop and returning to the coop as the first video acquisition strategy and the close-up video acquisition strategy as the second video acquisition strategy as an example, the process of detecting whether the recognition results meet the source conditions corresponding to the video acquisition strategy and the process of generating the source video are explained in detail.
[0115] (1) First video acquisition strategy
[0116] To efficiently and accurately obtain the source video corresponding to the first video acquisition strategy, in an optional implementation of this embodiment, when the video acquisition strategy corresponding to the video in the active area is the first video acquisition strategy, the following method can be used to detect whether the identification result meets the source tracing conditions corresponding to the video acquisition strategy:
[0117] Read the historical recognition results of the associated active area videos of the active area video;
[0118] If the historical identification results are inconsistent with the identification results, it is determined that the identification results meet the tracing conditions;
[0119] If the historical identification result is consistent with the identification result, it is determined that the identification result does not meet the tracing conditions, and no action is taken.
[0120] Optionally, the associated activity area video includes the activity area video acquired at a second acquisition time associated with the first acquisition time, determined based on the first acquisition time of the activity area video; for example, if the acquisition time of the activity area video is 8:00 AM, then the associated activity area video is the activity area video acquired at 7:00 AM. That is, the associated activity area video can be the activity area video acquired at the previous acquisition time corresponding to the acquisition time of the activity area video. For example, the associated acquisition time for the acquisition time 8:00 AM is 7:00 AM; the associated acquisition time for the acquisition time 7:50 AM is 7:40 AM. It should be noted that the previous acquisition time corresponding to the acquisition time of the activity area video includes the previous acquisition time in the corresponding video acquisition strategy.
[0121] Specifically, if the biometric behavior recognition results of the activity area video are inconsistent with the biometric behavior recognition results of the associated activity area video, the recognition results are determined to meet the tracing conditions corresponding to the video acquisition strategy.
[0122] For example, if the video capture strategy for the activity area video is "chickens leaving the coop," and the prompt text is "Are there chickens currently present?", and the biometric recognition of the activity area video is performed using the above method, and the recognition result is "yes," then based on the first capture time of the activity area video at 8:00 AM, the associated second capture time is determined to be 7:00 AM. The activity area video at 7:00 AM is then read as the associated activity area video. If the biometric recognition result for the associated activity area video is "no," then if chickens are present in the activity area video at 8:00 AM but not in the activity area video at 7:00 AM, it indicates that there was chicken movement from the coop to the activity area between 7:00 AM and 8:00 AM. Therefore, the recognition result for this activity area video meets the traceability conditions. It should be noted that if the biometric recognition result for the associated activity area video is "yes," it means that the chickens were in the activity area between 7:00 AM and 8:00 AM, and there was no chicken movement from the coop to the activity area. Therefore, the recognition result for this activity area video does not meet the traceability conditions.
[0123] Furthermore, if the identification result meets the tracing conditions corresponding to the video acquisition strategy, and the tracing video of the object is generated based on the video of the active area, in order to ensure the validity of the generated tracing video, in an optional implementation method provided in this embodiment, the process of generating the tracing video of the object based on the video of the active area can be implemented in the following way:
[0124] Read the first acquisition time of the video of the activity area and the second acquisition time of the associated video of the activity area, and determine the source tracing time interval based on the first acquisition time and the second acquisition time;
[0125] The activity area videos collected within the time interval of the source tracing are identified as the source tracing videos.
[0126] Specifically, since the associated activity area video is the activity area video collected at the previous collection time corresponding to the activity area video, the activity area video collected within the tracing time interval from the second collection time to the first collection time is determined as the tracing video; in other words, the activity area video collected within the tracing time interval formed by the collection time of the associated activity area video of the activity area video and the collection time of the activity area video is determined as the tracing video.
[0127] Following the previous example, the video of the activity area collected between 7:00 and 8:00 in the morning will be used as the source video under the chicken leaving the chicken coop video collection strategy.
[0128] That is, step S208 can also be replaced by, if the identification result satisfies the tracing conditions corresponding to the video acquisition strategy, then the tracing time interval is determined based on the first acquisition time and the associated second acquisition time of the video in the activity area, and the video in the activity area acquired within the tracing time interval is used as the tracing video of the object. One or more other steps provided in this embodiment can be selected to form a new implementation method.
[0129] To further improve the effectiveness of source tracing and avoid using long-duration activity area videos as source tracing videos, which would result in watching a large amount of invalid video during the source tracing process, a divide-and-conquer method can be used to generate source tracing videos. Specifically, in one optional implementation of this embodiment, when the identification result is inconsistent with the historical identification result, the following method can be used to determine whether the identification result meets the source tracing conditions:
[0130] Read the time index of the first acquisition time of the video in the activity area, and detect whether the time index is a preset index;
[0131] If so, then the identification result is determined to satisfy the tracing condition;
[0132] If not, the associated time interval is split to obtain at least one sub-time interval;
[0133] The source video of the object is generated based on the recognition results corresponding to the activity area video of each sub-time period.
[0134] Optionally, the associated time interval is determined based on the first acquisition time and the second acquisition time.
[0135] Specifically, when the recognition results of the activity area video and the associated activity area video are inconsistent, if the time index of the collection time of the activity area video is a preset index, then the recognition result is determined to meet the tracing conditions. Then, the associated time interval from the second collection time (associated collection time) to the first collection time (collection time) is taken as the tracing time interval, and the activity area video collected within the tracing time interval is taken as the tracing video. If the time index of the collection time of the activity area video is not a preset threshold, then the associated time interval is divided into sub-times to obtain at least one sub-time. Based on the recognition results of the activity area videos of adjacent sub-times in at least one sub-time, it is detected whether the recognition results of adjacent sub-times meet the corresponding tracing conditions. Then, the tracing video of the object is generated based on the activity area videos of adjacent sub-times.
[0136] It should be noted that the above process can be performed cyclically until the time index of the sub-time period reaches the preset index. Specifically, if the time index of the acquisition time of the activity area video does not reach the preset threshold, the associated time interval is divided into sub-time periods to obtain at least one sub-time period. The identification results of the activity area video at each sub-time period are checked against the identification results of the associated activity area video. If they are consistent, no further processing is required. If not, the time index of the first sub-time period (the acquisition time of the activity area video whose identification result is inconsistent with the associated activity area video) is checked against the preset index. If it is, the associated time interval from the second sub-time period (the sub-time period preceding the first sub-time period) to the first sub-time period is taken as the source time interval, and the activity area video acquired within the source time interval is taken as the source video. If not, the associated time interval from the second sub-time period to the first sub-time period is further divided into sub-time periods, and the above process is repeated. It should be noted that during the sub-time division process, sub-time division can be performed with priorities of 30 minutes, 10 minutes, 5 minutes, and 1 minute. The preset index can be the index corresponding to 1 minute.
[0137] In addition to ending the divide-and-conquer method based on preset indicators, the interval duration of the associated time interval can also be used to end the divide-and-conquer method. For example, if the recognition result of the activity area video is inconsistent with the recognition result of the associated activity area video, check whether the interval duration of the associated time interval formed by the second acquisition time and the first acquisition time is the preset duration, such as 10 minutes. If not, the associated time interval is divided into sub-times. If so, it is determined that the recognition result meets the traceability conditions.
[0138] Continuing with the example of chickens leaving the coop and moving to the activity area between 7:00 AM and 8:00 AM, if this behavior is detected (i.e., the 7:00 AM detection result is negative, and the 8:00 AM detection result is positive), since the associated time interval between 7:00 AM and 8:00 AM is one hour, if a chicken leaves the coop and moves to the activity area at 7:50 AM, the source video for the activity area from 7:00 AM to 7:50 AM will be empty. Directly using the activity area video from 7:00 AM to 8:00 AM as the source video would result in prolonged viewing of an empty activity area video, which would negatively impact... To improve the traceability experience, the time interval from 7:00 to 8:00 is divided into sub-time periods. First, it's divided into 30-minute intervals: 7:00, 7:30, and 8:00. The identification results of the activity area video at 7:00 (which can be video collected within that minute at 7:00) are checked against the identification results at 7:30, and then against the identification results at 7:30 and 8:00. If the identification results of the activity area video at 7:00 and 7:30 are consistent, then the traceability condition is not met, and no further action is taken. However, the recognition results of the activity area video at 7:30 AM are inconsistent with those at 8:00 AM. Therefore, it is checked whether the 30-minute interval between 7:30 AM and 8:00 AM is equal to the preset 10-minute interval. Since the interval length is not equal to the preset length, the associated time interval between 7:30 AM and 8:00 AM is further divided into 10-minute sub-times, resulting in 7:30 AM, 7:40 AM, 7:50 AM, and 8:00 AM. The second acquisition time associated with 7:30 AM is empty; the second acquisition time associated with 7:40 AM is 7:30 AM; the second acquisition time associated with 7:50 AM is 7:40 AM; and the second acquisition time associated with 8:00 AM is empty. The collection time is 7:50. The identification results of the activity area videos collected at each collection time are checked to see if they are consistent with the identification results of the associated second collection time. If they are consistent, it is determined that the traceability conditions are not met. After the test, the identification results of the activity area videos at 8:00 are inconsistent with the identification results of the activity area videos collected at the associated collection time of 7:50. Moreover, the duration of the associated time interval from 7:50 to 8:00 is a preset duration of 10 minutes. Therefore, the associated time interval from 7:50 to 8:00 is taken as the traceability time interval, and the activity area videos collected within this traceability time interval from 7:50 to 8:00 are taken as a traceability video corresponding to the video collection strategy of "chickens leaving the chicken coop".
[0139] Furthermore, when the prompt message of the chicken leaving the coop video acquisition strategy indicates whether the chicken leaving the coop behavior exists, if the identification result is yes, it is determined that the corresponding traceability condition is met; if the identification result is no, it is determined that the corresponding traceability condition is not met. That is, step S208 can also be replaced by generating a traceability video of the object based on the video of the identified area when the identification result is a positive identification result, and one or more other steps provided in this embodiment can be selected to form a new implementation method.
[0140] (2) Second video acquisition strategy
[0141] In another optional implementation of this embodiment, the following method can be used to detect whether the recognition result meets the tracing conditions corresponding to the video acquisition strategy: if the recognition result is not empty, determine that the recognition result meets the tracing conditions; if the recognition conditions are empty, determine that the recognition result does not meet the tracing conditions; correspondingly, if the recognition result is not empty, determine the video of the active area as the tracing video.
[0142] It should be noted that the different types of video acquisition strategies can be omitted, and the detection of whether the traceability conditions are met can be performed according to either of the two methods mentioned above. This embodiment does not limit this. In addition, step S208 can also be replaced by, if the identification result meets the traceability conditions corresponding to the video acquisition strategy, then the video of the active area is taken as the traceability video of the object, and one or more other steps provided in this embodiment can be selected to form a new implementation method.
[0143] In the specific execution process, after generating the traceability video of the object, video description information of the traceability video can also be generated based on the video acquisition strategy, recognition results and / or video information of the traceability video (such as video time); for example, a traceability video of chickens leaving the chicken coop at 7:50 am on x year x month x day, or a traceability video of chickens jumping at t hour on x year x month x day.
[0144] After generating the object's source tracing video, the video can be stored for object tracing. It should be noted that one video acquisition strategy can correspond to multiple source tracing videos. In one optional implementation provided in this embodiment, object tracing can be performed in the following manner:
[0145] Based on the user's request to trace the object, the source tracing video of the object is read and sent to the user.
[0146] It should also be noted that the video processing method provided in this embodiment can be implemented according to the configuration of the video acquisition strategy. For example, if the video acquisition strategy for chickens leaving the coop is configured to detect at every hour on the hour, then the above process will be executed at every hour on the hour. In addition, the video acquisition strategy for chickens leaving the coop can also be configured to detect at 10:00 AM, then detection will be performed starting from the video of the activity area at every hour on the hour at 10:00 AM. This embodiment does not limit this.
[0147] In summary, the video processing method provided in this embodiment configures a video acquisition component and a video acquisition strategy in the active area of the object, enabling the video acquisition component to acquire video according to the strategy without continuous video acquisition, thus avoiding wear and tear on the video acquisition component caused by acquiring invalid video. After acquiring the video of the active area of the object captured by the video acquisition component, the corresponding video acquisition strategy is determined based on the acquisition time and / or duration of the video in the active area, and the prompt information of the video acquisition strategy is read. Then, the video in the active area is encoded by a video encoder to obtain video features, and the prompt information is encoded by an information encoder to obtain information features. The video features and information features are then input into a large language model for biometric behavior recognition to obtain recognition results. In this way, the large language model is introduced to improve the convenience and efficiency of biometric behavior recognition. When the recognition results meet the traceability conditions corresponding to the video acquisition strategy, a traceability video of the object is generated based on the video in the active area, and the traceability video and video description information are stored so that they can be sent to the user during subsequent object traceability processes, improving the effectiveness and convenience of traceability and enhancing the user's perception of the object.
[0148] The following example illustrates the application of a video processing method provided in this embodiment in the generation of traceability videos for farmed animals. Figure 6 The video processing method provided in this embodiment for generating traceability videos of farmed animals will be further described below. Figure 6 A video processing method applied to the generation of traceability videos for farmed animals includes the following steps.
[0149] Step S602: Obtain video of the activity area captured by the camera configured in the activity area of the captive animals.
[0150] Step S604: Determine the corresponding video acquisition strategy based on the acquisition time of the video in the activity area, and read the prompt text of the video acquisition strategy.
[0151] Step S606: Perform video frame skipping on the video in the active area to obtain a video frame sequence.
[0152] In step S608, the video encoder performs video encoding on the video frame sequence to obtain video features.
[0153] In step S610, the text encoder performs text encoding on the prompt text to obtain text features.
[0154] The execution order of steps S610 and steps S606 to S608 is not limited here.
[0155] Step S612: Input the video features and text features into the model to perform biological behavior recognition and obtain the recognition results.
[0156] Step S614: Determine the videos of adjacent activity areas based on the acquisition time of the videos of the activity areas.
[0157] Optionally, adjacent active area videos include active area videos acquired at the previous acquisition time corresponding to the acquisition time of the active area video.
[0158] Step S616: Detect whether the recognition result of the active area video is consistent with the recognition result of the adjacent active area video;
[0159] If so, no action is needed;
[0160] If not, proceed to step S618.
[0161] Step S618: Detect whether the interval duration of the associated time interval between the second acquisition time of the adjacent active area video and the first acquisition time of the active area video is a preset duration.
[0162] If so, proceed to step S622.
[0163] If not, proceed to step S620.
[0164] Step S620: Divide the associated time interval into sub-time segments to obtain at least one sub-time segment, and return the video of the active area of each sub-time segment to execute steps S614 to S616.
[0165] Step S622: Use the video of the activity area collected within the associated time interval as the source video of the captive animals under the video collection strategy.
[0166] Subsequently, video description information can be generated based on the video time and video acquisition strategy of the source video, and the source video and video description information can be stored together.
[0167] Furthermore, steps S602 to S622 can be replaced by: when the detection condition of the video acquisition strategy is triggered, acquiring at least one video of the active area of the farmed animal at the acquisition time corresponding to the video acquisition strategy; the video encoder performs video encoding on the video frame sequence corresponding to each active area video to obtain video features; the text encoder performs text encoding on the prompt text of the video acquisition strategy to obtain text features; the text features and the video features of each active area video are input into a large language model for biological behavior recognition to obtain the recognition results of each active area video; the adjacent active area videos of each active area video are determined according to the acquisition time of each active area video; at least one target active area video whose recognition result is inconsistent with the recognition result of the adjacent active area video is selected, and the active area videos acquired from the second acquisition time (acquisition time of the adjacent active area video) of each target active area video to the first acquisition time of the target active area video are used as the source videos of the farmed animal.
[0168] It should be noted that any one or more steps in steps S602 to S622 can be combined with any one or more steps in steps S202 to S208 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features in steps S602 to S622 can be selected and combined with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S602 to S622 can be replaced with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.
[0169] The following example illustrates the application of a video processing method provided in this embodiment in the generation of traceability videos for farmed animals. Figure 7 The video processing method provided in this embodiment for generating traceability videos of farmed animals will be further described below. Figure 7 A video processing method applied to the generation of traceability videos for farmed animals includes the following steps.
[0170] Step S702: Obtain video of the activity area captured by the camera configured in the activity area of the captive animals.
[0171] Step S704: Determine the corresponding video acquisition strategy based on the acquisition duration of the video in the activity area, and read the prompt text of the video acquisition strategy.
[0172] Step S706: The video encoder performs video encoding on the video frame sequence corresponding to the active area video to obtain video features.
[0173] In step S708, the text encoder performs text encoding on the prompt text to obtain text features.
[0174] The execution order of steps S706 and S708 is not limited in this embodiment.
[0175] Step S710: Input video features and text features into the model for biological behavior recognition and obtain the recognition results.
[0176] Step S712: If the identification result is not empty, use the video of the active area as the source video of the captive animals.
[0177] Furthermore, steps S702 to S712 can be replaced by: when the detection condition of the video acquisition strategy is triggered, acquiring at least one video of the activity area of the captive animals at the acquisition time corresponding to the video acquisition strategy; performing video encoding on the video frame sequence corresponding to each activity area video through a video encoder to obtain video features, and performing text encoding on the prompt text of the video acquisition strategy through an information encoder to obtain text features; inputting the text features and the video features of each activity area video into a large language model for biological behavior recognition to obtain the recognition results of each activity area video, and using the activity area video with a non-empty recognition result as the source video of the captive animals.
[0178] It should be noted that any one or more steps in steps S702 to S712 can be combined with any one or more steps in steps S202 to S208 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features in steps S702 to S712 can be selected and combined with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S702 to S712 can be replaced with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.
[0179] This specification provides an embodiment of a video processing device as follows:
[0180] In the above embodiments, a video processing method is provided, and correspondingly, a video processing apparatus is also provided, which will be described below with reference to the accompanying drawings.
[0181] Reference Figure 8This diagram illustrates the structure of a video processing apparatus provided in this embodiment. Since the apparatus embodiment corresponds to the method embodiment, the description is relatively simple; relevant parts can be found in the corresponding descriptions of the method embodiments provided above. The apparatus embodiment described below is merely illustrative.
[0182] This embodiment provides a video processing apparatus, including:
[0183] The video acquisition module 802 is used to acquire the video of the active area of the object, determine the video acquisition strategy corresponding to the video of the active area, and read the prompt information of the video acquisition strategy;
[0184] Encoding module 804 is used for video encoder to encode the video of the active area to obtain video features, and information encoder to encode the prompt information to obtain information features;
[0185] The biometric behavior recognition module 806 is used to input the video features and the information features into the model to perform biometric behavior recognition and obtain recognition results.
[0186] If the identification result meets the tracing conditions corresponding to the video acquisition strategy, then the video generation module 808 is run. The video generation module 808 is used to generate a tracing video of the object based on the video of the activity area.
[0187] In one embodiment, when obtaining the information features, the following steps are performed:
[0188] The prompt information is positionally encoded to obtain positional features, and the prompt information is input into the embedding layer for information transformation to obtain initial information features;
[0189] The initial information features and the location features are fused, and the fused features are input into an autoregressive model for autoregressive calculation to obtain the information features.
[0190] In one embodiment, after the biometric behavior recognition module 806 inputs the video features and the information features into the model for biometric behavior recognition and obtains the recognition result, it further performs the following steps:
[0191] Read the historical recognition results of the associated active area videos of the active area video;
[0192] If the historical identification results are inconsistent with the identification results, it is determined that the identification results meet the tracing conditions.
[0193] In one embodiment, when the video generation module 808 generates a source video of the object based on the video of the active area, it performs the following steps:
[0194] Read the first acquisition time of the video of the activity area and the second acquisition time of the associated video of the activity area, and determine the source tracing time interval based on the first acquisition time and the second acquisition time;
[0195] The activity area videos collected within the time interval of the source tracing are identified as the source tracing videos.
[0196] In one embodiment, when the biometric behavior recognition module 806 determines that the recognition result meets the tracing conditions, it performs the following steps:
[0197] Read the time index of the first acquisition time of the video in the activity area, and detect whether the time index is a preset index;
[0198] If so, then the identification result is determined to satisfy the tracing condition;
[0199] If not, the associated time interval is split to obtain at least one sub-time interval, wherein the associated time interval is determined based on the first acquisition time and the second acquisition time;
[0200] The source video of the object is generated based on the recognition results corresponding to the activity area video of each sub-time period.
[0201] In one embodiment, after inputting the video features and the information features into the model for biometric behavior recognition and obtaining the recognition result, the biometric behavior recognition module 806 further performs the following steps:
[0202] If the identification result is not empty, it is determined that the identification result satisfies the tracing condition;
[0203] Accordingly, when generating the source video of the object based on the video of the active area, the video generation module 808 performs the following steps:
[0204] The video of the activity area is identified as the source video.
[0205] In one embodiment, before the video encoder encodes the video of the active region to obtain video features, the encoding module 804 further performs the following steps:
[0206] Video frames are extracted from the video of the active area to obtain a video frame sequence;
[0207] The video frame sequence is input into the video encoder.
[0208] In the video processing apparatus provided in this embodiment, during video processing, after acquiring the video of the active area of an object, the video acquisition strategy corresponding to the active area video is first determined. Then, the prompt information of the video acquisition strategy is read, thereby associating the active area video with the video acquisition strategy. The prompt information of the video acquisition strategy prompts subsequent biometric behavior recognition, improving the flexibility and effectiveness of biometric behavior recognition of the active area video. Based on this, the video encoder performs video encoding on the active area video to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features. Finally, the video features and information features are input into the model for biometric behavior recognition to obtain... Based on the identification results, before performing biometric behavior recognition through the model, the activity area video and prompt information are encoded separately to obtain accurate encoding results for the activity area video and prompt information. The encoded video features and information features are then used as input to enable the model to generate behavior recognition, thereby improving the accuracy of the model's biometric behavior recognition. After using the model to perform biometric behavior recognition on the activity area video corresponding to the video acquisition strategy and obtaining the recognition results, the source tracing conditions corresponding to the video acquisition strategy influence the generation of source tracing videos. Thus, based on the video acquisition strategy, biometric behavior recognition and source tracing video generation are performed, improving the richness and effectiveness of the generated source tracing videos.
[0209] This specification provides the following embodiment of a video processing device:
[0210] Corresponding to the video processing method described above, based on the same technical concept, this application also provides a video processing device for executing the video processing method provided above. Figure 9 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application.
[0211] This embodiment provides a video processing device, including:
[0212] like Figure 9As shown, video processing devices can vary significantly due to differences in configuration or performance. They may include one or more processors 901 and memory 902, with memory 902 storing one or more application programs or data. Memory 902 can be temporary or persistent storage. The application programs stored in memory 902 may include one or more modules (not shown), each module including a series of computer-executable instructions from the video processing device. Furthermore, processor 901 may be configured to communicate with memory 902, executing the series of computer-executable instructions stored in memory 902 on the video processing device. The video processing device may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input / output interfaces 905, one or more keyboards 906, etc.
[0213] In one specific embodiment, the video processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the video processing device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:
[0214] Acquire the video of the active area of the object, determine the video acquisition strategy corresponding to the video of the active area, and read the prompt information of the video acquisition strategy;
[0215] The video encoder performs video encoding on the video of the active area to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features;
[0216] The video features and information features are input into the model for biological behavior recognition to obtain the recognition result;
[0217] If the identification result satisfies the tracing conditions corresponding to the video acquisition strategy, a tracing video of the object is generated based on the video of the activity area.
[0218] In one embodiment, when the processor obtains the information features, it performs the following steps:
[0219] The prompt information is positionally encoded to obtain positional features, and the prompt information is input into the embedding layer for information transformation to obtain initial information features;
[0220] The initial information features and the location features are fused, and the fused features are input into an autoregressive model for autoregressive calculation to obtain the information features.
[0221] In one embodiment, after inputting the video features and the information features into the model for biometric behavior recognition and obtaining the recognition result, the processor further performs the following steps:
[0222] Read the historical recognition results of the associated active area videos of the active area video;
[0223] If the historical identification results are inconsistent with the identification results, it is determined that the identification results meet the tracing conditions.
[0224] In one embodiment, when the processor generates the source video of the object from the active region video, it performs the following steps:
[0225] Read the first acquisition time of the video of the activity area and the second acquisition time of the associated video of the activity area, and determine the source tracing time interval based on the first acquisition time and the second acquisition time;
[0226] The activity area videos collected within the time interval of the source tracing are identified as the source tracing videos.
[0227] In one embodiment, when the processor determines that the identification result meets the tracing conditions, it performs the following steps:
[0228] Read the time index of the first acquisition time of the video in the activity area, and detect whether the time index is a preset index;
[0229] If so, then the identification result is determined to satisfy the tracing condition;
[0230] If not, the associated time interval is split to obtain at least one sub-time interval, wherein the associated time interval is determined based on the first acquisition time and the second acquisition time;
[0231] The source video of the object is generated based on the recognition results corresponding to the activity area video of each sub-time period.
[0232] In one embodiment, after inputting the video features and the information features into the model for biometric behavior recognition and obtaining the recognition result, the processor further performs the following steps:
[0233] If the identification result is not empty, it is determined that the identification result satisfies the tracing condition;
[0234] Accordingly, in one embodiment, when the processor generates the source video of the object based on the active region video, it performs the following steps:
[0235] The video of the activity area is identified as the source video.
[0236] In one embodiment, before the video encoder encodes the video of the active region to obtain video characteristics, the processor further performs the following steps:
[0237] Video frames are extracted from the video of the active area to obtain a video frame sequence;
[0238] The video frame sequence is input into the video encoder.
[0239] In the video processing device provided in this embodiment, during video processing, after acquiring the video of the active area of an object, the video acquisition strategy corresponding to the active area video is first determined. Then, the prompt information of the video acquisition strategy is read, thereby associating the active area video with the video acquisition strategy. The prompt information of the video acquisition strategy prompts subsequent biometric behavior recognition, improving the flexibility and effectiveness of biometric behavior recognition of the active area video. Based on this, the video encoder performs video encoding on the active area video to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features. Finally, the video features and information features are input into the model for biometric behavior recognition to obtain... Based on the identification results, before performing biometric behavior recognition through the model, the activity area video and prompt information are encoded separately to obtain accurate encoding results for the activity area video and prompt information. The encoded video features and information features are then used as input to enable the model to generate behavior recognition, thereby improving the accuracy of the model's biometric behavior recognition. After using the model to perform biometric behavior recognition on the activity area video corresponding to the video acquisition strategy and obtaining the recognition results, the source tracing conditions corresponding to the video acquisition strategy influence the generation of source tracing videos. Thus, based on the video acquisition strategy, biometric behavior recognition and source tracing video generation are performed, improving the richness and effectiveness of the generated source tracing videos.
[0240] This specification provides an embodiment of a computer-readable storage medium as follows:
[0241] Corresponding to the video processing method described above, based on the same technical concept, this application also provides a computer-readable storage medium.
[0242] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process:
[0243] Acquire the video of the active area of the object, determine the video acquisition strategy corresponding to the video of the active area, and read the prompt information of the video acquisition strategy;
[0244] The video encoder performs video encoding on the video of the active area to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features;
[0245] The video features and information features are input into the model for biological behavior recognition to obtain the recognition result;
[0246] If the identification result satisfies the tracing conditions corresponding to the video acquisition strategy, a tracing video of the object is generated based on the video of the activity area.
[0247] In the computer-readable storage medium provided in this embodiment, during video processing, after acquiring the video of the active area of the object, the video acquisition strategy corresponding to the active area video is first determined. Then, the prompt information of the video acquisition strategy is read, thereby associating the active area video with the video acquisition strategy. The prompt information of the video acquisition strategy prompts subsequent biometric behavior recognition, improving the flexibility and effectiveness of biometric behavior recognition of the active area video. Based on this, the video encoder performs video encoding on the active area video to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features. Finally, the video features and information features are input into the model for biometric behavior recognition. To obtain the recognition results, before performing biometric behavior recognition through the model, the activity area video and prompt information are encoded separately to obtain accurate encoding results. The encoded video features and information features are then used as input to enable the model to generate behavior recognition, thereby improving the accuracy of the model's biometric behavior recognition. After obtaining the recognition results, the source tracing conditions corresponding to the video acquisition strategy influence the generation of source tracing videos. Thus, based on the video acquisition strategy, biometric behavior recognition and source tracing video generation are performed, improving the richness and effectiveness of the generated source tracing videos.
[0248] It should be noted that the embodiments of a computer-readable storage medium described in this specification and the embodiments of a video processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0249] Another embodiment of this disclosure also provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the following process:
[0250] Acquire the video of the active area of the object, determine the video acquisition strategy corresponding to the video of the active area, and read the prompt information of the video acquisition strategy;
[0251] The video encoder performs video encoding on the video of the active area to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features;
[0252] The video features and information features are input into the model for biological behavior recognition to obtain the recognition result;
[0253] If the identification result satisfies the tracing conditions corresponding to the video acquisition strategy, a tracing video of the object is generated based on the video of the activity area.
[0254] In the computer program product provided in this embodiment, during video processing, after acquiring the video of the active area of the object, the video acquisition strategy corresponding to the active area video is first determined. Then, the prompt information of the video acquisition strategy is read, thereby associating the active area video with the video acquisition strategy. The prompt information of the video acquisition strategy prompts subsequent biometric behavior recognition, improving the flexibility and effectiveness of biometric behavior recognition of the active area video. Based on this, the video encoder performs video encoding on the active area video to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features. Finally, the video features and information features are input into the model for biometric behavior recognition. To obtain the identification results, before performing biometric behavior identification through the model, the activity area video and prompt information are encoded separately to obtain accurate encoding results. The encoded video features and information features are then used as input to enable the model to generate behavior identification, thereby improving the accuracy of the model's biometric behavior identification. After obtaining the identification results, the biometric behavior identification corresponding to the video acquisition strategy is performed on the activity area video corresponding to the video acquisition strategy. The source tracing conditions corresponding to the video acquisition strategy influence the generation of source tracing videos. Thus, biometric behavior identification and source tracing video generation are performed based on the video acquisition strategy, improving the richness and effectiveness of the generated source tracing videos.
[0255] The computer program product in this disclosure embodiment can implement the various processes of the above-described video processing method embodiment and achieve the same effects and functions, which will not be repeated here.
[0256] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0257] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, this specification can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0258] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable test apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable test apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0259] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable test equipment to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0260] These computer program instructions can also be loaded onto a computer or other programmable test equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0261] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0262] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0263] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0264] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0265] The embodiments of this application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0266] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0267] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.
Claims
1. A video processing method, characterized in that, The method includes: Acquire video of the activity area of the object, determine the video acquisition strategy corresponding to the video of the activity area, and read the prompt information of the video acquisition strategy; the video acquisition strategy includes video acquisition time, video acquisition duration, prompt information and / or tracing conditions, and the prompt information is used to instruct the model to perform biological behavior recognition on the video of the activity area in accordance with the prompt information; The video encoder performs video encoding on the video of the active area to obtain video features, and the information encoder performs information encoding on the prompt information to obtain information features; The video features and the information features are input into the model to perform biological behavior recognition and obtain the recognition result; If the identification result satisfies the tracing conditions corresponding to the video acquisition strategy, a tracing video of the object is generated based on the activity area video; wherein, the tracing conditions include the inconsistency between the identification result and the historical identification result of the associated activity area video of the activity area video, and the associated activity area video includes the activity area video acquired at the previous acquisition time corresponding to the acquisition time of the activity area video.
2. The method according to claim 1, characterized in that, The acquired information features include: The prompt information is position-encoded to obtain position features, and the prompt information is initial information-encoded to obtain initial information features; The initial information features and the location features are fused, and the fused features are input into an autoregressive model for autoregressive calculation to obtain the information features.
3. The method according to claim 1, characterized in that, After the step of inputting the video features and the information features into the model for biometric behavior recognition and obtaining the recognition result is executed, the method further includes: Read the historical recognition results of the associated active area videos of the active area video; the associated active area videos include the active area videos collected at the previous collection time corresponding to the collection time of the active area video; If the historical identification results are inconsistent with the identification results, it is determined that the identification results meet the tracing conditions.
4. The method according to claim 3, characterized in that, The process of generating the source video of the object based on the video of the active area includes: Read the first acquisition time of the video of the activity area and the second acquisition time of the associated video of the activity area, and determine the source tracing time interval based on the first acquisition time and the second acquisition time; The activity area videos collected within the time interval of the source tracing are identified as the source tracing videos.
5. The method according to claim 4, characterized in that, The step of determining that the identification result satisfies the tracing condition includes: Read the time index of the first acquisition time of the video in the activity area, and detect whether the time index is a preset index; If so, then the identification result is determined to satisfy the tracing condition; If not, the associated time interval is split to obtain at least one sub-time interval, wherein the associated time interval is determined based on the first acquisition time and the second acquisition time; The source video of the object is generated based on the recognition results corresponding to the activity area video of each sub-time period.
6. The method according to claim 1, characterized in that, After the step of inputting the video features and the information features into the model for biometric behavior recognition and obtaining the recognition result is executed, the method further includes: If the identification result is not empty, it is determined that the identification result satisfies the tracing condition; The process of generating the source video of the object based on the video of the active area includes: The video of the activity area is identified as the source video.
7. The method according to claim 1, characterized in that, Before the video encoder performs video encoding on the active region video to obtain video features, the process further includes: Video frames are extracted from the video of the active area to obtain a video frame sequence; The video frame sequence is input into the video encoder.
8. A video processing apparatus, characterized in that, The device includes: The video acquisition module is used to acquire video of the activity area of the object, determine the video acquisition strategy corresponding to the video of the activity area, and read the prompt information of the video acquisition strategy; the video acquisition strategy includes video acquisition time, video acquisition duration, prompt information and / or tracing conditions, and the prompt information is used to instruct the model to perform biological behavior recognition on the video of the activity area in accordance with the prompt information. The encoding module is used for the video encoder to encode the video of the active area to obtain video features, and the information encoder to encode the prompt information to obtain information features. A biometric behavior recognition module is used to input the video features and the information features into the model to perform biometric behavior recognition and obtain recognition results; If the identification result satisfies the tracing conditions corresponding to the video acquisition strategy, then the video generation module is run. The video generation module is used to generate a tracing video of the object based on the activity area video. The tracing conditions include the inconsistency between the identification result and the historical identification result of the associated activity area video. The associated activity area video includes the activity area video acquired at the previous acquisition time corresponding to the acquisition time of the activity area video.
9. A video processing device, characterized in that, The device includes: A processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to perform the video processing method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store computer-executable instructions that, when executed by a processor, implement the video processing method as described in any one of claims 1-7.