Image Recognition Method, Apparatus, Electronic Device, and Readable Storage Medium

By obtaining and combining video frame image feature information, the video status is automatically determined, and the problem of artificial detection of "food broadcast" illegal videos and high live broadcast costs is solved, and efficient and accurate illegal detection is achieved.

CN114092860BActive Publication Date: 2025-07-29VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111416865.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-07-29
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

In the prior art, the detection cost of artificially detecting illegal videos and live broadcasts of "food broadcasts" is relatively high.

Method used

By obtaining frame image feature information of the target video, including portrait, food, sequence, time and image features, combining adjacent frame images to form a clip, and automatically determine the video status based on the clip feature information, and using the target model for detection.

Benefits of technology

It realizes automatic, accurate and efficient detection of whether the video is in a violation state, reducing the detection cost of human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092860B_ABST
    Figure CN114092860B_ABST
Patent Text Reader

Abstract

The present application discloses an image recognition method, apparatus, electronic device, and readable storage medium, belonging to the technical field of image processing. Among them, the method includes: obtaining the feature information of N1 frame images of a target video, where N1 is a positive integer, and the feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information; combining at least two adjacent frame images with matching feature information to obtain N2 first segments of the target video, where N2 is a positive integer and N2 < N1; obtaining the combined feature information of the first segment based on the feature information of each frame image in the combined first segment; determining the first state of the target video based on the combined feature information of the N2 first segments; where the first state is any one of a violation state and a non-violation state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and specifically relates to an image recognition method, device, electronic device and readable storage medium. Background Art

[0002] In recent years, the internet short video, livestreaming, and influencer economy have flourished, but some of the attendant irregularities have also surfaced. For example, some "eating broadcast" bloggers, driven by a desire for traffic, engage in excessive eating, not only turning food into a "gluttonous feast" but also running counter to the ideals of healthy eating and food conservation. Consequently, major platforms are cracking down on this phenomenon, detecting illegal "eating broadcast" videos and livestreams and imposing appropriate penalties.

[0003] Existing technologies require manual strategies to detect illegal "eating broadcast" videos and live streams through human judgment. For example, they can manually determine whether a video is a normal food sharing activity or an unhealthy eating habit; another example is to manually classify the severity of unhealthy eating broadcasts and impose different penalties.

[0004] It can be seen that in the existing technology, manual detection of illegal videos and live broadcasts such as "eating broadcasts" leads to high detection costs. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide an image recognition method that can solve the problem of high detection costs caused by manual detection of illegal videos and live broadcasts of the "eating broadcast" type in the prior art.

[0006] In a first aspect, an embodiment of the present application provides an image recognition method, the method comprising: obtaining feature information of N1 frame images of a target video, where N1 is a positive integer, and the feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information; combining at least two adjacent frame images with matching feature information, and obtaining N2 first segments of the target video, where N2 is a positive integer, and N2 <N1;基于组合所述第一片段的各个帧图像的特征信息,获取所述第一片段的组合特征信息;基于N2个所述第一片段的组合特征信息,确定所述目标视频的第一状态;其中,所述第一状态为违规状态和非违规状态中的任一项。

[0007] Second aspect, an embodiment of the present application provides an image recognition device, which includes: a first acquisition module, configured to acquire feature information of N1 frame images of a target video, where N1 is a positive integer, and the feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information; a combination module, configured to combine at least two adjacent frame images with matching feature information, and obtain N2 first segments of the target video, where N2 is a positive integer, and N2 < N1; a second acquisition module, configured to acquire combined feature information of the first segment based on the feature information of each frame image in the combined first segment; a determination module, configured to determine a first state of the target video based on the combined feature information of the N2 first segments; where the first state is any one of a violation state and a non-violation state.

[0008] Third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0009] Fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0010] Fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0011] In this way, in the embodiment of the present application, first, N1 frame images of the target video are acquired, and the feature information of each frame image is acquired respectively. The feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information. Secondly, adjacent frame images with matching feature information are combined together to form N2 first segments. The matching methods include, but are not limited to, feature information with the same scene or the same object. Further, the feature information of all frame images in the first segment is fused to obtain the combined feature information of the first segment. Finally, based on the combined feature information of all first segments, the first state of the target video is determined, that is, it is determined whether the target video belongs to a violation state or a non-violation state. It can be seen that in the embodiment of the present application, by acquiring and combining the feature information of the video, it is possible to automatically, accurately, and efficiently determine whether the video belongs to a violation state, so as to achieve the purpose of detecting violation behaviors in the video, and avoid human intervention, with a relatively low detection cost. Description of the Drawings

[0012] Figure 1 is a flowchart of the image recognition method according to an embodiment of the present application;

[0013] Figures 2 to 5 is an explanatory schematic diagram of the image recognition method according to an embodiment of the present application;

[0014] Figure 6 is a block diagram of the image recognition device according to an embodiment of the present application;

[0015] Figure 7 is one of the schematic diagrams of the hardware structure of the electronic device according to an embodiment of the present application;

[0016] Figure 8 is the second schematic diagram of the hardware structure of the electronic device according to an embodiment of the present application. Detailed implementation manners

[0017] Next, the technical solutions of the embodiments of the present application will be clearly described in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0018] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0019] Next, the image recognition method provided by the embodiments of the present application will be described in detail in conjunction with the accompanying drawings, through specific embodiments and their application scenarios.

[0020] Figure 1 shows a flowchart of the image recognition method according to an embodiment of the present application. This method is applied to an electronic device and includes:

[0021] Step 110: Obtain the feature information of N1 frame images of the target video, where N1 is a positive integer, and the feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information.

[0022] Optionally, the image recognition method in this embodiment is used for detecting violations in videos, live broadcasts, etc.

[0023] Optionally, the target video is a short video uploaded to a certain platform; the target video is a live broadcast on a certain platform.

[0024] Optionally, the duration corresponding to one frame image is: one second in the target video.

[0025] Optionally, N1 frame images of the target video are obtained at a preset frequency.

[0026] Optionally, the frequency is set moderately. On the one hand, to avoid too long an interval between the extracted frame images due to too low a frequency, which is not sufficient to express the entire video content; on the other hand, to avoid too many learning objects due to too high a frequency.

[0027] Thus, in this step, a set of frame images of the target video can be obtained.

[0028] Furthermore, the feature information of each frame image is obtained separately.

[0029] Optionally, the violation detection in this embodiment is: the detection of bad eating behaviors. Therefore, the feature information obtained in this step is related to eating behaviors, including portrait feature information and food feature information; in addition, the obtained frame images are used to reflect the entire video content, so the feature information also includes sequence feature information and time feature information; furthermore, the feature information also includes image feature information.

[0030] Optionally, the image recognition method of this embodiment can be implemented through a target model. Therefore, the obtained feature information needs to be input into the target model, and the following is a detailed description of the input process for each item of feature information.

[0031] First of all, it should be noted that generally, if there is someone eating continuously for a long time in the video and there is a lot of food, it is considered that there is a bad eating behavior in the video.

[0032] Therefore, first, obtain portrait feature information to assist in explaining whether someone is eating.

[0033] Among them, the portrait feature information is encoded to be converted into a feature form acceptable to the target model.

[0034]

[0035]

[0036] Table 1

[0037] For example, referring to Table 1, if there is no portrait, the portrait feature information is "0", and if there is a portrait, the portrait feature information is "1".

[0038] Second, obtain food feature information to assist in explaining whether someone is eating.

[0039] Among them, encode the food feature information to convert it into a feature form acceptable to the target model.

[0040] Food characteristics Digital coding Absent food 0 Present food 1

[0041] Table 2

[0042] Referring to Table 2, if there is no food, the food feature information is "0"; if there is food, the food feature information is "1".

[0043] Furthermore, obtaining the food feature information is also used to assist in explaining the size of the food.

[0044] Generally, the size of the food is an important feature for distinguishing between bad eating behaviors and normal eating behaviors. In bad eating behaviors, there is a large amount of food to attract attention. Therefore, if the area occupied by the food in the image is large, it indicates that there may be someone eating a large amount, such as overeating, which belongs to bad eating behaviors.

[0045] For reference, obtain the proportion of the food area in the image, and use the proportion of the food area in the image as the food feature information.

[0046] Optionally, encode the food feature information to convert it into a feature form acceptable to the target model.

[0047]

[0048] Table 3

[0049] Referring to Table 3, if there is no food, the food feature information is "0"; if the proportion of the food area in the image is less than 10%, the food is small, and the food feature information is "1"; if the proportion of the food area in the image is less than 30%, the food is medium, and the food feature information is "2"; if the proportion of the food area in the image is greater than 30%, the food is large, and the food feature information is "3".

[0050] For reference, locate the food in the image to obtain the area occupied by the food, and use the ratio of this area to the entire image area as the proportion of the food area in the image. Among them, the upper left corner of the image can be selected as the coordinate origin to establish a coordinate system, find the coordinates of each vertex of the area occupied by the food, and thus calculate the proportion of the food area in the image.

[0051] Third, obtain sequence feature information to assist in explaining the duration of eating in the entire video; in addition, sort the obtained frame images according to this sequence for subsequent grouping and restoring the order of behaviors in the video, etc.

[0052] Fourth, obtain time feature information to represent the occurrence time of the eating behavior, so as to assist in explaining the duration of eating in the entire video.

[0053] The time feature information includes the playing time point of the frame image in the target video, such as minutes and seconds.

[0054] Among them, the time feature information and the sequence feature information are encoded to be converted into a feature form acceptable to the target model.

[0055] Remarks Digital coding 0:00:01 The first second 1 0:00:02 The second second 2 0:00:03 The third second 3 0:00:04 The fourth second 4

[0056] Table 4

[0057] Referring to Table 4, the time feature information includes: "0:00:01", "0:00:02", "0:00:03", "0:00:04"; the sequence feature information includes: "0", "1", "2", "3", "4".

[0058] Fifth, obtain image feature information to restore the image features of the frame image itself.

[0059] For example, the image feature information includes the picture presented by a certain frame image.

[0060] Optionally, use object detection technology (such as using the selected yolov5 model) to locate the people and food in the frame image, and record the time point of the frame image. Further, locate the food to obtain a positioning box, and record the upper left point x and the lower right point y of the box, and calculate the proportion of the food area in the picture.

[0061] It should be noted that the yolo series of object detection models can identify and locate the objects in the image by reading the image once, which greatly improves the object recognition efficiency. Yolov5 is iterated on the basis of the series and has a faster recognition speed, which is a relatively advanced object detection technology.

[0062]

[0063]

[0064] Table 5

[0065] Referring to Table 5, in this embodiment, the feature information of each frame image is obtained.

[0066] Further, map each item of feature information into a feature vector acceptable to the target model, corresponding to:

[0067] fea food-size = emd (degree of food size); corresponding to the feature vector used to represent the degree of food size;

[0068] feafood = emd(Whether there is food); corresponding to the feature vector used to represent whether there is food;

[0069] fea peopel = emd(Whether there is a human figure); corresponding to the feature vector used to represent whether there is a human figure.

[0070] Furthermore, these feature vectors are concatenated to obtain the constructed additional feature vector:

[0071] fea extra = concat( food-size

fea food

[0072] In addition, the frame image itself is encoded through the backbone network Efficientnetv2 to obtain an image feature vector:

[0073] fea pic = Efficientnetv2(frame image).

[0074] It should be noted that Efficientnetv2 is a convolutional neural network. Compared with other convolutional network models, it has a faster training speed and better parameter efficiency. As a basic backbone network for extracting image features, it has good performance in downstream tasks such as image classification.

[0075] Furthermore, fea position = emd(time point); corresponding to the feature vector used to represent the time point; among them, the feature vector used to represent the sequence can be obtained from fea position , and it can also be understood as the feature vector representing the position.

[0076] Optionally, this embodiment can use the transfomer encoder for encoding.

[0077] It should be noted that the transfomer is a neural network model that uses the self-attention mechanism to encode sequence features. The parallelized design and the complexity of the model make the accuracy and performance of this model higher than those of the previously popular RNN recurrent neural network.

[0078] Among them, the encoding part of the transfomer is not limited to encoding the feature information of the frame image in this embodiment, but can also encode the combined feature information of subsequent segments to form a feature form that can be received by the target model.

[0079] Further, input the feature vectors of each frame image obtained above into the target model.

[0080] It should be noted that for the processes of obtaining and encoding feature information in this embodiment, reference can be made to Figure 2 as shown.

[0081] Step 120: Combine at least two adjacent frame images with matching feature information to obtain N2 first segments of the target video, where N2 is a positive integer and N2 < N1.

[0082] In this step, combine the N1 frame images and use the combined first segment as the detection unit for detection, which can reduce the number of learning objects.

[0083] Preferably, the combination rule is as follows: First, for the N1 frame images obtained, arrange them in the chronological order from front to back in the target video. Second, in each frame image, find the frame images with matching feature information. Finally, combine the frame images with matching and adjacent feature information together.

[0084] For example, for a travel sharing video with a duration of two minutes, the first minute is food sharing and the second minute is scenery sharing. Therefore, based on all the frame images obtained in the first minute, all of them include the feature information of food, and these frame images are combined together; based on all the frame images obtained in the second minute, none of them include the feature information of food, and these frame images are combined together.

[0085] Step 130: Obtain the combined feature information of the first segment based on the feature information of each frame image in the combined first segment.

[0086] Preferably, refer to Figure 3 , construct the feature vector fea of any number of frame images extra , the image feature vector fea pic , the time series feature vector fea position , after adding and combining them, use it as the feature vector of one frame image; further, based on any first segment, use the feature vectors of the frame images used to combine this first segment as the input (the feature vector used to represent the input is: fea block-input = fea extra + fea pic + fea position ), and input them into the target model respectively. After the feature vectors of multiple frame images are fused by the transfomer encoder, output the feature vector of the first segment (the feature vector used to represent the output is: fea block-output = transfomer(fea block-input))), and the feature vector of the first segment is used to represent the combined feature information of the first segment.

[0087] Among them, the first segment corresponds to a video block in the illustration.

[0088] Step 140: Determine the first state of the target video based on the combined feature information of N2 first segments.

[0089] Among them, the first state is either a violation state or a non-violation state.

[0090] In this step, after fusing the feature information of all frame images in each first segment, the combined feature information corresponding to the first segment can be obtained, and thus, based on the combined feature information of N2 first segments, the first state of the target video is determined.

[0091] Referentially, if it is analyzed from the combined feature information of N2 first segments that there is a long-term continuous eating behavior in the target video and the food area involved is large, then the violation state of the target video is determined; otherwise, the non-violation state of the target video is determined.

[0092] In this way, in the embodiment of the present application, first, N1 frame images of the target video are obtained, and the feature information of each frame image is respectively obtained. The feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information. Secondly, adjacent frame images with matching feature information are combined together to form N2 first segments. The matching methods include, but are not limited to, having the same scene or the same object feature information. Further, the feature information of all frame images in the first segment is fused to obtain the combined feature information of the first segment. Finally, based on the combined feature information of all first segments, the first state of the target video is determined, that is, it is determined whether the target video belongs to the violation state or the non-violation state. It can be seen that in the embodiment of the present application, by obtaining, combining, etc. the feature information of the video, it is possible to automatically, accurately, and efficiently determine whether the video belongs to the violation state, so as to achieve the purpose of detecting violation behaviors in the video, and avoid manual intervention, with a relatively low detection cost.

[0093] In addition, the graph recognition method in the present application is implemented based on a target model. Therefore, during the training process of the target model, the state can be determined for the training samples based on the above steps in the embodiment. Among them, the training samples can be a database including a large number of videos, and for each video in the database, its corresponding state, that is, the violation state or the non-violation state, has been determined based on other relatively accurate methods (such as manual judgment). Thus, based on the states of each video obtained by the target model, the target model can be trained so that the probability that the states of a large number of videos obtained by the target model match their actual states reaches a certain threshold.

[0094] In the process of the image recognition method according to another embodiment of the present application, step 120 includes:

[0095] Sub-step A1: Combine at least two adjacent frame images that all include portrait feature information and do not include food feature information into a second segment.

[0096] Sub-step A2: Combine at least two adjacent frame images that all include portrait feature information and include food feature information into a third segment.

[0097] Sub-step A3: Combine at least two adjacent frame images that do not include portrait feature information and do not include food feature information into a fourth segment.

[0098] Sub-step A4: Combine the second segment, the third segment, and the fourth segment to obtain N2 first segments.

[0099] Generally, in a video, the behaviors of the captured people are continuous. Therefore, in two frame images with a short interval duration, there may be the same objects, scenes, etc. Thus, based on this phenomenon, in this embodiment, N1 frame images are combined into N2 first segments.

[0100] Exemplarily, in the detection of whether there is an improper eating behavior in a video, the video scene is divided into: a single-person scene, a single-food scene, a person-and-food scene, and a no-person-and-no-food scene, so that adjacent frame images in the same scene can be merged.

[0101] For reference, first, if a series of consecutive frame images obtained all include portrait feature information and do not include food feature information, they are combined into a second segment. The second segment corresponds to a single-person scene, and there is no eating behavior in the second segment.

[0102] Among them, one second segment corresponds to one video block.

[0103] Second, if a series of consecutive frame images obtained all include portrait feature information and include food feature information, they are combined into a third segment. The third segment corresponds to a person-and-food scene, and there is an eating behavior in the third segment.

[0104] Among them, one third segment corresponds to one video block.

[0105] Third, if a series of consecutive frame images obtained do not include portrait feature information and do not include food feature information, they are combined into a fourth segment. The fourth segment corresponds to a no-person-and-no-food scene, and there is no eating behavior in the fourth segment.

[0106] Among them, a fourth segment corresponds to a video block.

[0107] See Figure 4 , for example, after combining each frame image of the target video, video blocks 1, 2, and 3 are obtained. Correspondingly, all the frame images in video block 1 include people and food, all the frame images in video block 2 include people, and all the frame images in video block 3 do not include people or food.

[0108] Furthermore, all the obtained second segments, third segments, and fourth segments form N2 first segments.

[0109] In this embodiment, some frame images with similar characteristics are combined as video blocks. In the same video block, all the images are similar and usually have certain identical semantics. For example, in a food sharing video, it includes the following video blocks: preparing food, cooking, close-up, eating, etc. For a violation video, there are often only video blocks such as eating, resting, eating, resting, etc., repeating continuously. Therefore, this embodiment performs unified processing in the form of video blocks. While retaining the key scene information in the video, it can also convert a large number of frame images into a small number of video blocks, which can greatly reduce the number of learning objects.

[0110] In addition, in combination with step 130 in this embodiment, through the combined feature information of each first segment, the relationship between each first segment (such as the repeated cycle of eating and resting) can be reflected, so as to detect the bad eating behavior existing in the video.

[0111] In the process of the image recognition method in another embodiment of the present application, step 140 includes:

[0112] Sub-step B1: Combine N3 first segments among the N2 first segments whose relevance satisfies the first preset condition, and obtain N4 fifth segments of the target video, where N3 and N4 are positive integers, and N4 < N2.

[0113] Sub-step B2: Combine N5 fifth segments among the N4 fifth segments whose relevance satisfies the second preset condition, and obtain N6 sixth segments of the target video, where N5 and N6 are positive integers, and N6 < N4, N3 < N5.

[0114] In this embodiment, after obtaining the N2 first segments, further, in one combination process, the N2 first segments are combined again, such as pairwise combination or three-way combination, etc., to obtain a smaller number of fifth segments; further, in another combination, the N4 fifth segments are combined again, such as pairwise combination or three-way combination, etc., to obtain a smaller number of sixth segments.

[0115] Among them, a fifth segment corresponds to a video block; a sixth segment corresponds to a video block.

[0116] Optionally, the number of combinations is not limited and may be related to the value of N2. The larger N2 is, the more times of combination, so as to gradually reduce the number of video blocks to be processed.

[0117] Optionally, the relevance satisfies the first preset condition, corresponding to: adjacent relationship; the relevance satisfies the second preset condition, corresponding to: adjacent relationship.

[0118] Among them, in the previous combination, the number of combined video blocks is: N3; in the next combination, the number of combined video blocks is: N5; N3 < N5. Based on such a limitation, the features of each adjacent video block can be fused in a hierarchical and progressive manner, realizing the fusion process of the pyramid structure from fine-grained to coarse-grained. Thus, through such a fusion process, for each finally obtained video block, its combined feature information can fully reflect the relationship between each video block, which helps to analyze whether there is an improper eating behavior.

[0119] Sub-step B3: Determine the first state of the target video based on the combined feature information of N6 sixth segments.

[0120] Optionally, through the target model, the combined feature information of N6 sixth segments is summed and averaged to determine the first state of the target video according to the obtained average value.

[0121] For reference, the average value of the features here can reflect the proportion of the duration of the portrait's eating behavior for large-area food in the video duration obtained by combining the scene information in the video.

[0122] Among them, if the scene information is not considered, through a single calculation, the proportions of the duration of the portrait's eating behavior for large-area food in the two videos are the same, and they can be determined as illegal videos or non-illegal videos at the same time. However, the proportion obtained based on this embodiment combines the scene information. For example, in the first video, the combined scene information is that the portrait eats intermittently, while in the second video, the combined scene is that the portrait eats continuously. Obviously, the probability of an illegal behavior in the second video is higher.

[0123] Among them, the scene information here can be reflected by the feature information of this application.

[0124] In this embodiment, in order to better explore the relationship between video blocks, the feature information of each adjacent video block is fused in a hierarchical and progressive manner, realizing the fusion process of the pyramid structure from fine-grained to coarse-grained. Finally, the overall feature information of the video is obtained, that is, the target model outputs the final video feature vector, so as to detect whether there is any illegal behavior in the video according to the video feature vector. It can be seen that based on such a method, the relationship between video blocks can be well analyzed, thereby improving the accuracy of detecting illegal videos.

[0125] In the process of the image recognition method according to another embodiment of the present application, step 140 includes:

[0126] Sub-step C1: Combine two adjacent first segments into a fifth segment.

[0127] See Figure 5 , before this step, construct the feature vector of the first segment. Among them, the constructed feature vector of the first segment is merged from the following several vectors:

[0128] Sequence feature vector: fea block-index =Emd(video block serial number), which is used to represent the position information of the first segment in the target video, such as which first segment;

[0129] Duration feature vector: fea block-time =Emd(video block duration), which is used to represent the duration of the first segment in the target video, such as the number of frames of the image extracted from the first segment in the target video, or the playing duration obtained according to the frame extraction frequency; this feature vector also represents the importance degree of the first segment in the target video;

[0130] The fea block-output (video block feature) obtained in the foregoing embodiment.

[0131] Furthermore, fuse the three of them as the constructed feature vector of the first segment:

[0132] Fea block =fea block-index +fea block-time +fea block-output .

[0133] Video block Video block sequence Video block duration The first block 0 5 The second block 1 4 … … …

[0134] Table 6

[0135] Among them, as shown in Table 6, some item feature information for constructing the feature vector of the first segment is shown, so as to map each item of feature information into a feature vector acceptable to the target model, that is, the constructed feature vector of the first segment.

[0136] SeeFigure 5 , the process of constructing the feature vector of the first segment corresponds to: the one-gram video block combination process.

[0137] Among them, in the target model, the one-gram video block combination process corresponds to:

[0138] One-gram = MulitHead(Fea block )。

[0139] Furthermore, based on the feature vectors of each constructed first segment, every two adjacent first segments are combined according to the serial number to form several fifth segments, which corresponds to the bigram video block combination process.

[0140] In the target model, the combination process of the bigram video block corresponds to:

[0141] Two-gram = MulitHead(concat(One-gram)).

[0142] It should be noted that multi-head attention (MulitHead) is a self-attention calculation method, which was first applied to the Transformer structural block. Compared with the traditional self-attention calculation, the calculated feature vectors are divided into multiple heads, and each head performs separate attention calculations, reducing the computational amount while enabling each head to have complex and independent semantic feature information.

[0143] Sub-step C2: Combine three adjacent fifth segments into one sixth segment.

[0144] Referring to the previous step, the obtained fifth segments are combined every three adjacent fifth segments according to the serial number to form several sixth segments, which corresponds to the trigram video block combination process.

[0145] Furthermore, in the target model, the combination process of the trigram video block corresponds to:

[0146] Three-gram = MulitHead(concat(Two-gram))

[0147] In this embodiment, the trinary progressive combination method is adopted. On the one hand, it achieves a hierarchical progressive combination effect, and on the other hand, it can effectively control the number of learning objects.

[0148] Sub-step C3: Determine the first state of the target video based on the combined feature information of each sixth segment.

[0149] Optionally, in the target model, the features of each sixth segment are output, summed and averaged to obtain the final representation vector output of the target video, and then passed through a fully connected layer to calculate the vector to obtain the final video prediction result pred.

[0150] Among them, output = sum(Three-gram) / size; pred = MLP(output).

[0151] Optionally, using cross-entropy loss to optimize the network can achieve the purpose of training the target model.

[0152] In this embodiment, a hierarchical video block fusion method is proposed. First, video block features are constructed; then, based on the video block features, fusion is performed from level 1 to level 3 to obtain the final video output features. Specifically, this embodiment adopts the n-gram pyramid feature fusion technology, which is fused and combined level by level. First, a unary input, that is, a single video block feature (finer), is used and fused using a multi-head self-attention network. Then, the network outputs are pairwise concatenated with adjacent video blocks to obtain a binary group, which is fused using a multi-head self-attention network. Finally, the obtained results are concatenated and output. Finally, a ternary feature combination is used to concatenate the features of three adjacent video blocks (coarser), and a multi-head attention network is used to output the feature vector. By the principle of from fine to coarse, the relationships between different video blocks are gradually fused, starting from a small range of unary groups to a large range of ternary groups, and finally the video is detected for features. It can be seen that based on this embodiment, the video block features can be well modeled, and the related video features can be fused to improve the accuracy of detection.

[0153] In the process of the image recognition method according to another embodiment of the present application, step 140 includes:

[0154] Sub-step D1: When an abnormal eating behavior of a person in the target video is detected based on the combined feature information of N2 first segments, determine the first state of the target video.

[0155] Among them, the first state is a violation state.

[0156] Optionally, the abnormal eating behavior of a person at least includes the following characteristics: there is a person eating, the eating duration is long, the continuity of the eating behavior is strong, and the area of the food in the eating behavior is large.

[0157] Optionally, based on the combined feature information of N2 first segments, through the target model, relevant numerical values can be obtained as the output result, which is used to measure the above characteristics, and then the video's belonging state is determined by comparing the output result with the preset conditions.

[0158] Correspondingly, when the output result meets the preset conditions, the video includes all of the above features, that is, there is abnormal eating behavior of a person in the video; otherwise, when the output result does not meet the preset conditions, the video does not include all of the above features, that is, there is no abnormal eating behavior of a person in the video.

[0159] In this embodiment, based on the analysis of the combined feature information of N2 first segments, it can be determined whether there is abnormal eating behavior of a person in the target video, and when there is abnormal eating behavior of a person in the target video, the result that the target video is in the first state is output, thereby completing the automatic detection process.

[0160] In summary, for the problem that it is difficult to automatically determine the scale of food-eating live broadcasts in videos, on the one hand, this application adds additional features of people and food, as well as time and sequence features to effectively reduce the dependence of the video understanding depth model on a large amount of training data; on the other hand, images with certain similar semantic information are combined into video blocks to reduce the difficulty of modeling a large number of images in a long video model and retain the key scene information in the video; on the third hand, based on the proposed video coding model, combined with the hierarchical progressive fusion process of n-gram, the combination of video block features is realized, so as to obtain the abnormal eating features of the video and effectively identify the abnormal eating behavior in the video.

[0161] It should be noted that for the image recognition method provided in the embodiment of this application, the execution subject may be an image recognition device, or a control module of the image recognition device for executing the image recognition method. In the embodiment of this application, the case where the image recognition device executes the image recognition method is taken as an example to illustrate the image recognition device provided in the embodiment of this application.

[0162] Figure 6 The block diagram of an image recognition device according to another embodiment of this application is shown, and the device includes:

[0163] A first acquisition module 10, configured to acquire the feature information of N1 frame images of a target video, where N1 is a positive integer, and the feature information includes human portrait feature information, food feature information, sequence feature information, time feature information, and image feature information;

[0164] A combination module 20, configured to combine at least two adjacent frame images with matching feature information, and obtain N2 first segments of the target video, where N2 is a positive integer, and N2 < N1;

[0165] A second acquisition module 30, configured to acquire the combined feature information of the first segment based on the feature information of each frame image of the combined first segment;

[0166] A determination module 40, configured to determine the first state of the target video based on the combined feature information of N2 first segments;

[0167] Wherein, the first state is either a violation state or a non-violation state.

[0168] Thus, in the embodiments of the present application, first, N1 frame images of the target video are obtained, and the feature information of each frame image is obtained respectively. The feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information. Secondly, adjacent frame images with matching feature information are combined together to form N2 first segments. The matching methods include, but are not limited to, feature information with the same scene or the same object. Further, the feature information of all the frame images in the first segment is fused to obtain the combined feature information of the first segment. Finally, based on the combined feature information of all the first segments, the first state of the target video is determined, that is, it is determined whether the target video belongs to the violation state or the non-violation state. It can be seen that in the embodiments of the present application, by obtaining and combining the feature information of the video, it is possible to automatically, accurately, and efficiently determine whether the video belongs to the violation state, so as to achieve the purpose of detecting violation behaviors in the video, and avoid human intervention with relatively low detection costs.

[0169] Optionally, the combination module 20 includes:

[0170] The first combination unit is used to combine at least two adjacent frame images that both include portrait feature information and do not include food feature information into a second segment;

[0171] The second combination unit is used to combine at least two adjacent frame images that both include portrait feature information and include food feature information into a third segment;

[0172] The third combination unit is used to combine at least two adjacent frame images that neither include portrait feature information nor include food feature information into a fourth segment;

[0173] The fourth combination unit is used to combine the second segment, the third segment, and the fourth segment to obtain N2 first segments.

[0174] Optionally, the determination module 40 includes:

[0175] The fifth combination unit is used to combine N3 first segments among the N2 first segments whose relevance meets the first preset condition, and obtain N4 fifth segments of the target video. N3 and N4 are positive integers, and N4 < N2;

[0176] The sixth combination unit is used to combine N5 fifth segments among the N4 fifth segments whose relevance meets the second preset condition, and obtain N6 sixth segments of the target video. N5 and N6 are positive integers, and N6 < N4, N3 < N5;

[0177] The first determination unit is configured to determine a first state of the target video based on the combined feature information of N6 sixth segments.

[0178] Optionally, the determination module 40 includes:

[0179] The seventh combination unit is configured to combine two adjacent first segments into one fifth segment;

[0180] The eighth combination unit is configured to combine three adjacent fifth segments into one sixth segment;

[0181] The second determination unit is configured to determine a first state of the target video based on the combined feature information of each sixth segment.

[0182] Optionally, the determination module 40 includes:

[0183] The third determination unit is configured to determine a first state of the target video when abnormal eating behavior of a person in the target video is detected based on the combined feature information of N2 first segments;

[0184] Wherein, the first state is a violation state.

[0185] The image recognition device according to the embodiment of the present application may be a device, or a component of a terminal, an integrated circuit, or a chip. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc., which are not specifically limited in the embodiment of the present application.

[0186] The image recognition device according to the embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0187] The image recognition device provided by the embodiment of the present application can implement each process implemented by the above method embodiment. To avoid repetition, it will not be elaborated here.

[0188] Optionally, asFigure 7 As shown in the figure, an embodiment of the present application further provides an electronic device 100, including a processor 101, a memory 102, a program or instruction stored on the memory 102 and executable on the processor 101. When the program or instruction is executed by the processor 101, it implements each process of any of the above image recognition method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0189] It should be noted that the electronic device in the embodiment of the present application includes the above-mentioned mobile electronic device and non-mobile electronic device.

[0190] Figure 8 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0191] The electronic device 1000 includes, but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010 and other components.

[0192] Those skilled in the art can understand that the electronic device 1000 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1010 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 8 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0193] Among them, the processor 1010 is used to obtain the feature information of N1 frame images of the target video, where N1 is a positive integer, and the feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information; combine at least two adjacent frame images that match the feature information, and obtain N2 first segments of the target video, where N2 is a positive integer, and N2 < N1; based on the feature information of each frame image of the combined first segment, obtain the combined feature information of the first segment; based on the combined feature information of the N2 first segments, determine the first state of the target video; where the first state is any one of a violation state and a non-violation state.

[0194] Thus, in the embodiments of the present application, first, N1 frame images of the target video are acquired, and the feature information of each frame image is acquired respectively. The feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information. Secondly, adjacent frame images with matching feature information are combined together to form N2 first segments. The matching methods include but are not limited to feature information with the same scene or the same object. Further, the feature information of all the frame images in the first segment is fused to obtain the combined feature information of the first segment. Finally, based on the combined feature information of all the first segments, the first state of the target video is determined, that is, it is determined whether the target video belongs to a violation state or a non-violation state. It can be seen that in the embodiments of the present application, by acquiring, combining, etc. the feature information of the video, it is possible to automatically, accurately, and efficiently determine whether the video belongs to a violation state, so as to achieve the purpose of detecting violation behaviors in the video, and avoid human intervention with relatively low detection costs.

[0195] Optionally, the processor 1010 is further configured to combine at least two adjacent frame images that both include portrait feature information and do not include food feature information into a second segment; combine at least two adjacent frame images that both include portrait feature information and include food feature information into a third segment; combine at least two adjacent frame images that neither include portrait feature information nor include food feature information into a fourth segment; and combine the second segment, the third segment, and the fourth segment to obtain N2 of the first segments.

[0196] Optionally, the processor 1010 is further configured to combine N3 of the first segments among the N2 first segments whose relevance satisfies a first preset condition, and obtain N4 fifth segments of the target video, where N3 and N4 are positive integers, and N4 < N2; combine N5 of the fifth segments among the N4 fifth segments whose relevance satisfies a second preset condition, and obtain N6 sixth segments of the target video, where N5 and N6 are positive integers, and N6 < N4, N3 < N5; and determine the first state of the target video based on the combined feature information of the N6 sixth segments.

[0197] Optionally, the processor 1010 is further configured to combine two adjacent first segments into a fifth segment; combine three adjacent fifth segments into a sixth segment; and determine the first state of the target video based on the combined feature information of each sixth segment.

[0198] Optionally, the processor 1010 is further configured to determine the first state of the target video when an abnormal eating behavior of a person in the target video is detected based on the combined feature information of the N2 first segments; where the first state is the violation state.

[0199] In summary, in view of the problem that it is difficult to automatically determine the scale of food-eating live broadcasts in videos, on the one hand, this application adds additional features of people and food, as well as time and sequence features, to effectively reduce the dependence of the video understanding deep model on a large amount of training data; on the other hand, images with certain similar semantic information are merged into video chunks, reducing the difficulty of modeling a large number of images in a long-video model and retaining the key scene information in the video; on the third hand, based on the proposed video coding model, combined with the hierarchical progressive fusion process of n-gram, the merging of video chunk features is realized, so as to obtain the bad eating features of the video, effectively identifying bad eating behaviors in the video.

[0200] It should be understood that in the embodiments of this application, the input unit 1004 may include a Graphics Processing Unit (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes the image data of static pictures or video images obtained by an image capture device (such as a camera) in the video image capture mode or the image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here. The memory 1009 may be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 1010 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1010.

[0201] The embodiments of this application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned image recognition method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0202] Among them, the processor is the processor of the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0203] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiment of the image recognition method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0204] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0206] The embodiments of the present application have been described above with reference to the drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. An image recognition method, characterized in that, The method includes: Obtaining the feature information of N1 frame images of the target video, where N1 is a positive integer, and the feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information; Combining at least two adjacent frame images with matching feature information, and obtaining N2 first segments of the target video, where N2 is a positive integer and N2 < N1; Based on the feature information of each frame image that composes the first segment, obtaining the combined feature information of the first segment; Based on the combined feature information of the N2 first segments, determining the first state of the target video; Wherein, the first state is any one of a violation state and a non-violation state; The combining at least two adjacent frame images with matching feature information and obtaining N2 first segments of the target video includes: Combining at least two adjacent frame images that both include portrait feature information and do not include food feature information into a second segment; Combining at least two adjacent frame images that both include portrait feature information and include food feature information into a third segment; Combining at least two adjacent frame images that neither include portrait feature information nor include food feature information into a fourth segment; Combining the second segment, the third segment, and the fourth segment to obtain N2 first segments.

2. The method according to claim 1, wherein The determining the first state of the target video based on the combined feature information of the N2 first segments includes: Combining N3 first segments among the N2 first segments whose relevance satisfies a first preset condition, and obtaining N4 fifth segments of the target video, where N3 and N4 are positive integers and N4 < N2; Combining N5 fifth segments among the N4 fifth segments whose relevance satisfies a second preset condition, and obtaining N6 sixth segments of the target video, where N5 and N6 are positive integers and N6 < N4, N3 < N5; Based on the combined feature information of the N6 sixth segments, determining the first state of the target video.

3. The method according to claim 1, wherein The determining the first state of the target video based on the combined feature information of the N2 first segments includes: Combining two adjacent first segments into a fifth segment; Combining three adjacent fifth segments into a sixth segment; Based on the combined feature information of each sixth segment, determining the first state of the target video.

4. The method according to claim 1, wherein The determining the first state of the target video based on the combined feature information of the N2 first segments includes: When detecting an abnormal eating behavior of a person in the target video based on the combined feature information of the N2 first segments, determining the first state of the target video; Wherein, the first state is the violation state.

5. An image recognition device, characterized in that, The device includes: A first acquisition module, configured to acquire the feature information of N1 frame images of the target video, where N1 is a positive integer, and the feature information includes portrait feature information, food feature information, sequence feature information, time feature information, and image feature information; A combination module, configured to combine at least two adjacent frame images with matching feature information, and obtain N2 first segments of the target video, where N2 is a positive integer and N2 < N1; A second acquisition module, configured to acquire combined feature information of the first segment based on the feature information of each frame image of the combined first segment; A determination module, configured to determine a first state of the target video based on the combined feature information of the N2 first segments; Wherein, the first state is any one of a violation state and a non-violation state; The combination module includes: A first combination unit, configured to combine at least two adjacent frame images that both include portrait feature information and do not include food feature information into a second segment; A second combination unit, configured to combine at least two adjacent frame images that both include portrait feature information and include food feature information into a third segment; A third combination unit, configured to combine at least two adjacent frame images that neither include portrait feature information nor include food feature information into a fourth segment; A fourth combination unit, configured to combine the second segment, the third segment, and the fourth segment to obtain N2 first segments.

6. The device according to claim 5, characterized in that The determination module includes: A fifth combination unit, configured to combine N3 first segments among the N2 first segments whose relevance satisfies a first preset condition, and obtain N4 fifth segments of the target video, where N3 and N4 are positive integers and N4 < N2; A sixth combination unit, configured to combine N5 fifth segments among the N4 fifth segments whose relevance satisfies a second preset condition, and obtain N6 sixth segments of the target video, where N5 and N6 are positive integers and N6 < N4, N3 < N5; A first determination unit, configured to determine the first state of the target video based on the combined feature information of the N6 sixth segments.

7. The device according to claim 5, wherein The determination module includes: A seventh combination unit, configured to combine two adjacent first segments into a fifth segment; An eighth combination unit, configured to combine three adjacent fifth segments into a sixth segment; A second determination unit, configured to determine the first state of the target video based on the combined feature information of each sixth segment.

8. The device according to claim 5, characterized in that, The determination module includes: A third determination unit, configured to determine the first state of the target video when abnormal eating behavior of a person in the target video is detected based on the combined feature information of the N2 first segments; Wherein, the first state is the violation state.

9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the image recognition method according to any one of claims 1 to 4 are implemented.

10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the image recognition method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Live broadcast mask generation method, readable storage medium and computer equipment

    CN112492323A

  • Video processing method and device and computer readable storage medium

    CN113449824A