A cross-modal livestock behavior retrieval method and device

By enhancing content and time semantics of livestock videos and texts, extracting multi-channel features and hierarchical fusion, the problem of inaccurate livestock behavior retrieval results is solved, and efficient livestock behavior monitoring and management is achieved.

CN119597968BActive Publication Date: 2025-08-12AGRI INFORMATION INST OF CHINESE ACAD OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411573078.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2025-08-12
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

In the prior art, the modal semantic consistency of livestock videos and text descriptions is poor, resulting in low accuracy of livestock behavior search results and it is difficult to quickly locate livestock individuals with specific behaviors.

Method used

By enhancing content semantics and temporal semantics for livestock videos and texts, multi-channel features of videos and text are extracted, synonyms collection and hierarchical fusion technology are used to improve semantic consistency between modals and achieve accurate retrieval.

Benefits of technology

Improve the accuracy of livestock video-text retrieval, reduce artificial control costs, and realize accurate and intelligent monitoring and management of large-scale livestock behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119597968B_ABST
    Figure CN119597968B_ABST
Patent Text Reader

Abstract

This invention provides a cross-modal livestock behavior retrieval method and device, relating to the field of text-video retrieval technology. The method comprises: obtaining a livestock video and user-entered input text, the input text including livestock behavior events; extracting a target sentence, nouns, and verbs from the input text; determining the target sentence's noun and verb retrieval tags based on a set of words corresponding to livestock behavior retrieval tags; searching for video features from a video-text multi-channel feature set that match the noun and verb retrieval tags and the sentence structure of the target sentence; and retrieving livestock behavior events from the livestock video based on the matched video features. This approach improves the accuracy of livestock behavior event retrieval from livestock videos based on the input text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text-video retrieval, and in particular to a cross-modal livestock behavior retrieval method and device. Background Art

[0002] With the development of image retrieval technology, image retrieval technology has gradually begun to be applied in the field of animal husbandry, so the accuracy of image retrieval is crucial.

[0003] Animal behavior generally reflects their physiological health. Promptly identifying abnormal animal behavior can prevent disease transmission, adjust feeding practices, and monitor breeding. Currently, farmers use cameras to record animal videos and images. However, these videos and images are rich in content and contain a large amount of livestock information. If farmers need to quickly retrieve video segments of livestock behaviors they are interested in from the vast amount of livestock monitoring videos and locate specific livestock individuals exhibiting specific behaviors, manual queries, while highly accurate, are inefficient.

[0004] Traditional methods can learn a shared embedding space to establish a connection between video and text. The main processing method is to directly retrieve the target video based on the input text. Due to the diversity of text descriptions input by different users and the poor semantic consistency between different modalities, the current retrieval results are inaccurate. For example, searching for a cow drinking water on the ground may result in a video of the cow standing by the water. Summary of the Invention

[0005] The purpose of the present invention is to address the problem of poor accuracy of retrieval results in the prior art and to provide a cross-modal livestock behavior retrieval method and device, which can improve the accuracy of retrieving livestock of target livestock behavior events from livestock videos based on input text.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a cross-modal livestock behavior retrieval method, which obtains livestock videos and input text input by a user, wherein the input text includes livestock behavior events; extracts target sentences, nouns, and verbs in the input text; determines the retrieval tags of nouns and verbs in the target sentence based on the word set corresponding to the livestock behavior retrieval tags; searches for video features that match the retrieval tags of nouns, verbs, and sentence structure of the target sentence from a video-text multi-channel feature set; retrieves livestock behavior events in the livestock video based on the matched video features; wherein the word set corresponding to the livestock behavior retrieval tags includes: a synonym set of the noun retrieval tags, and a synonym set of the verb retrieval tags; the video-text multi-channel features are fused features after temporal semantic enhancement and content semantic enhancement are performed on both the descriptive text and the video corresponding to the livestock behavior events.

[0008] In a second aspect, an embodiment of the present invention provides a cross-modal livestock behavior retrieval device, which includes: a data acquisition module, a feature extraction module, a retrieval label determination module, a feature matching module and a retrieval module; the data acquisition module is used to acquire livestock videos and input text input by a user, the input text including livestock behavior events; the feature extraction module is used to extract target sentences, nouns and verbs in the input text; the retrieval label determination module is used to determine the retrieval labels of nouns and the retrieval labels of verbs in the target sentence based on the word set corresponding to the livestock behavior retrieval label; the feature extraction module is used to search for video features that match the retrieval labels of nouns, the retrieval labels of verbs and the sentence structure of the target sentence from a video-text multi-channel feature set; the feature matching module is used to retrieve livestock behavior events in livestock videos based on the matched video features; wherein the word set corresponding to the livestock behavior retrieval label includes: a synonym set of the noun retrieval label and a synonym set of the verb retrieval label; the video-text multi-channel feature is a fusion feature after the description text and video corresponding to the livestock behavior event are both subjected to temporal semantic enhancement and content semantic enhancement.

[0009] According to a third aspect of an embodiment of the present invention, a computer device is provided, comprising: a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any of the above cross-modal livestock behavior retrieval methods are implemented.

[0010] According to a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above cross-modal livestock behavior retrieval methods are implemented.

[0011] The cross-modal livestock behavior retrieval method provided by the embodiments of the present invention, firstly, enriches the essential semantic information of livestock videos by performing content semantic enhancement and temporal semantic enhancement on livestock videos. It also enriches the essential semantic information of livestock input text by performing content semantic enhancement and temporal semantic enhancement on livestock input text, thereby obtaining essential semantic information about the entire event under different modalities. Secondly, based on the features of the semantically enhanced livestock videos and the features of the semantically enhanced input text, they are layered and fused into multi-channel video features and multi-channel text features. This layered fusion combines the complementary features of the two different modalities, video and text, reducing redundant information in the fused features, obtaining a more comprehensive understanding of the entire livestock behavior event, and enhancing semantic consistency between different modalities. This method can effectively improve the accuracy of livestock video-text retrieval, reduce the cost of manual control, and achieve precise and intelligent real-time monitoring and management of large amounts of livestock behavior. For example, based on the input text entered by the user, livestock behavior videos of interest can be accurately and quickly retrieved from a large amount of agricultural livestock monitoring videos, and livestock with abnormal behavior can be located. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a schematic diagram of a model training process for the cross-modal livestock behavior retrieval method provided in an embodiment of the present invention.

[0013] Figure 2 The second schematic diagram of the model training process of the cross-modal livestock behavior retrieval method provided in an embodiment of the present invention.

[0014] Figure 3 A semantically enhanced logic diagram provided by an embodiment of the present invention.

[0015] Figure 4 A schematic diagram of a cross-modal livestock behavior retrieval training model provided by an embodiment of the present invention.

[0016] Figure 5 A schematic diagram of a cross-modal livestock behavior retrieval training logic provided by an embodiment of the present invention.

[0017] Figure 6 A schematic flow chart of a cross-modal livestock behavior retrieval method provided by an embodiment of the present invention.

[0018] Figure 7 A hardware structure diagram of a computer device housing a cross-modal livestock behavior retrieval system provided in an embodiment of the present invention.

[0019] Figure 8 A schematic diagram of the structure of a cross-modal livestock behavior retrieval device is provided for an embodiment of the present invention. DETAILED DESCRIPTION

[0020] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0021] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0022] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."

[0023] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0024] First, the model training process of the cross-modal livestock retrieval method provided by an embodiment of the present invention is given.

[0025] Figure 1 A schematic diagram of a model training process of a cross-modal livestock behavior retrieval method provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown in , the method includes the following S101 to S105:

[0026] S101. Obtain a training video and a description text matching the livestock behavior events in the training video.

[0027] The training video is any video in the training video set.

[0028] Specifically, livestock behavior events usually refer to specific behaviors or activities related to farm animals or livestock (such as cattle, sheep, pigs, etc.). Livestock behavior events can be livestock behaviors that last for a period of time, livestock behaviors that last for a long time, or livestock behaviors that last for a short time, such as feeding behavior, breeding behavior, social behavior, alertness and escape behavior, migration and foraging behavior, disease or discomfort behavior, training and obedience behavior, etc. The embodiments of the present invention do not specifically limit this.

[0029] It is understood that the description text is a manually input text description that matches the livestock behavior events in the training video. It is understood that the description texts input by different users are diverse and different.

[0030] For example, for a cow ruminating, user 1 enters the description text "a cow ruminating," user 2 enters the description text "a cow chewing for a long time without biting any grass," user 3 enters the description text "the cow is ruminating," and user 4 enters the description text "the cow is chewing all the time."

[0031] S102: Perform content semantic enhancement on the training video and the description text, and perform time semantic enhancement on the training video and the description text to obtain semantically enhanced video features and semantically enhanced text features.

[0032] Typically, the content semantics of livestock behavior events involve the behavior patterns, habits, and interactions of livestock (such as cows, sheep, pigs, etc.).

[0033] For example, the content semantics describing the livestock behavior event in the training video may include at least one of the following information: the type of livestock object where the livestock behavior event occurred, the specific object individual, the location where the event occurred, and the specific livestock behavior action.

[0034] It can be understood that the temporal semantics of livestock behavior events involve the behavior patterns and habits of livestock within a specific time period, as well as their relationship with the environment. For example, they can describe behavior timing, seasonal changes, growth and development stages, environmental influences, event sequences, and temporal patterns.

[0035] Exemplarily, the temporal semantics describing livestock behavior events in the training video may include at least one of the following information: the start time of the livestock behavior event (starting video frame), the duration of the behavior event (the continuous video frames during which the behavior event lasts), and the end time of the behavior event (ending video frame).

[0036] For example, the content semantics describing the livestock behavior event in the description text may include at least one of the following information: a sentence description of the livestock behavior event, a livestock object description, a livestock action description, and a location description.

[0037] For example, the temporal semantics of a livestock behavior event described in a description text may include at least one of the following information: tense information of the verb of the livestock behavior event, time information contained in the verb, and time information contained in the noun.

[0038] It can be understood that in the embodiment of the present invention, cross-modal semantic enhancement can be performed, that is, both the training video and the description text are semantically enhanced. Cross-modal semantic enhancement can obtain the essential semantic information of the entire event of livestock behavior.

[0039] By enhancing the temporal semantics of livestock behavior events in training videos, we can obtain the essential semantic information of livestock behavior throughout time.

[0040] For example, the essential semantics of cattle rumination behavior is that food in the cow's stomach returns to the cow's mouth through vomiting and is chewed again. This behavior can last for a period of time. Therefore, by enhancing the temporal semantics of livestock behavior events in training videos, the rumination behavior of cattle can be accurately described instead of chewing grass.

[0041] Specifically, content semantic enhancement and temporal semantic enhancement are performed on both the training video and the text, thereby improving the consistency between the content and key temporal information semantics of livestock behavior events presented in the training video and described in the text.

[0042] S103. Based on a predefined fusion loss function, the semantically enhanced video features and the semantically enhanced text features are layered-fused to obtain multi-channel features of the video-text pair after feature fusion.

[0043] For example, layer-by-layer fusion can be performed through a CNN network based on a residual attention mechanism.

[0044] S104: Perform text-video retrieval training based on a predefined text-video similarity function, a predefined multimodal retrieval loss function, and multi-channel features of video-text pairs.

[0045] S105. Iteratively train based on the text-video retrieval results to optimize the fusion loss function, similarity function, and multimodal retrieval loss function.

[0046] Specifically, based on the accuracy of the retrieval results, the coefficients and hyperparameters of the fusion loss function, the similarity coefficient of the text-video similarity function, and the loss coefficients and hyperparameters of the multimodal retrieval loss function can be optimized, so that the subsequent text-training video retrieval tasks based on the learned fusion loss function, similarity function and multimodal retrieval loss function can be more accurate.

[0047] An embodiment of the present invention provides a model training method for a cross-modal livestock behavior retrieval method. First, a training video and a descriptive text matching the livestock behavior events in the training video are obtained, content semantic enhancement is performed on the training video and the descriptive text, and time semantic enhancement is performed on the training video and the descriptive text to obtain semantically enhanced video features and semantically enhanced text features. Then, based on a predefined fusion loss function, the semantically enhanced video features and the semantically enhanced text features are layered fused to obtain feature-fused video-text pair multi-channel features. Thereafter, text-video retrieval training is performed based on a predefined text-video similarity function, a predefined multimodal retrieval loss function, and the video-text pair multi-channel features. Finally, iterative training is performed based on the text-video retrieval results to optimize the fusion loss function, the similarity function, and the multimodal retrieval loss function. First, by performing content and temporal semantic enhancement on training videos, the essential semantic information of livestock behavioral events in the videos can be enriched. By performing content and temporal semantic enhancement on the descriptive text, the essential semantic information of the descriptive text can also be enriched, thereby obtaining essential semantic information about the entire livestock behavioral event across different modalities. Second, based on the semantically enhanced video features and text features describing the livestock behavioral event, hierarchical feature fusion is performed to form multi-channel features for video-text pairs. This hierarchical feature fusion combines the complementary features of the two different modalities, video and text, reducing redundant information in the fused features, achieving a more comprehensive understanding of the entire livestock behavioral event, and enhancing semantic consistency across different modalities. This can effectively improve the accuracy of livestock video-text retrieval, reduce manual control costs, and enable precise and intelligent real-time monitoring and management of large-scale livestock behaviors. For example, based on user-entered input text, relevant livestock behavior videos can be accurately and quickly retrieved from massive amounts of agricultural livestock surveillance videos, allowing the location of livestock exhibiting abnormal behavior.

[0048] Optionally, in an embodiment of the present invention, the content semantic enhancement process in S102 may include the following steps A01 to A02:

[0049] A01. Enhance the local features and global features of the training videos describing livestock behavior events in the training videos.

[0050] It can be understood that local features in the training video can represent livestock behavior events from various presentation angles, while global features in the training video can represent livestock behavior events from the overall presentation angle.

[0051] For example, local features describing livestock behavior events in a training video include object features and motion features; global features describing livestock behavior events in a livestock training video include appearance features of the training video.

[0052] Among them, object features can represent livestock entities (such as cows) in livestock behavior events, motion features represent the movement of livestock entities in livestock behavior events (drinking water), and training appearance features can roughly describe the overall information of livestock behavior events (cows standing on the grass and drinking water).

[0053] A02. Enhance the local and global features of livestock behavior events in the description text.

[0054] It can be understood that the local features corresponding to the text indicate the characteristics of a specific part or fragment in the text, and can characterize livestock behavior events from multiple descriptive angles. For example, the local features describing livestock behavior events in the text include nouns and verbs. The global features corresponding to the text describe various indicators and information of the overall characteristics and attributes of the entire text, and can comprehensively describe livestock behavior events. For example, the global features describing livestock behavior events in the text include sentences. The sentences in the text can show the overall description of the event. In the embodiment of the present invention, sentences are used to represent the global semantic information of livestock behavior events, and the nouns and verbs in the sentences are used to understand the entities and the behaviors of the entities in the sentences.

[0055] Based on this scheme, by semantically enhancing both the local features and the global features describing livestock behavior events in the training video, the semantic information of the local features and the semantic information of the global features of the livestock training video can be enhanced respectively. By semantically enhancing both the local features and the global features describing livestock behavior events in the text, the semantic information of the local features and the semantic information of the global features of the text can be enhanced respectively, enriching the overall semantic information of the livestock behavior events and the semantic information of each detail, thereby enhancing the semantic consistency of the livestock training video and the descriptive text, thereby improving the accuracy of subsequent retrieval of training video segments or images in the livestock training video containing livestock that conform to the livestock behavior described in the descriptive text based on the semantically enhanced features.

[0056] Optionally, in the embodiment of the present invention, the semantic enhancement process of the training video content in A01 may include the following A11. Figure 2 As shown in , the content semantic enhancement process of the text in A02 may include the following A12 to A14:

[0057] A11. Extract object features, motion features, and appearance features of training videos.

[0058] It should be noted that the appearance feature can roughly describe the global information of the entire livestock behavior event in the training video, the cow drinking water, the object feature describes the local information of the training video of the livestock behavior event, the object is the cow, and the motion feature describes the local information of the training video of the livestock behavior event, the action is drinking water.

[0059] In an embodiment of the present invention, livestock behavior events are represented by extracting multi-channel features of training videos, that is, features of three training video channels, namely object features, motion features and appearance features, can be extracted from the training videos to characterize livestock behavior events from multiple angles.

[0060] For example, based on the object classification network, the object features of each frame sampled in the training video can be extracted to obtain a sequence of object features represents the deep feature vector of the oth frame in the i-th training video extracted by the object classification network. Where i is an integer from 1 to th, th represents the total number of training videos, and th is a positive integer.

[0061] For example, the behavior-based video retrieval network can extract the motion features of each frame sampled in the training video to obtain a sequence of motion features represents the deep features of the oth frame in the i-th training video extracted by the action video retrieval network.

[0062] For example, the pre-trained CNN (Convolutional Neural Networks) can be used to extract the training appearance features of each frame sampled in the th training video to obtain a sequence of appearance features: represents the deep features of the kth frame in the i-th training video extracted by the pre-trained CNN network.

[0063] For example, the sequence of appearance features is transformed into Aggregate into vector representation Obtain the global features of the training video; aggregate the object feature sequence of the training video into a vector representation through the NetVLAD model And the motion feature sequence of the training video Aggregated into vector representation as Get the local features of the training video.

[0064] It can be understood that by extracting semantic features describing livestock behavioral events in training videos in a multi-channel form, compared with related technologies, more fine-grained features can be extracted from training videos, making livestock behavioral events more distinguishable. That is, extracting appearance features, object features, and motion features from training videos can more richly describe the essential semantics of livestock behavioral events and improve the distinguishability of features. For example, rumination is a continuous action, including gastric flow and chewing, and generally 40s-50s training video segments can be retrieved, while related technologies may only be able to find training videos of chewing for a few seconds, and the retrieval accuracy is low.

[0065] A12. Obtain the description sentences in the description text, and the nouns and verbs in the description sentences.

[0066] A13. Extract synonyms that describe nouns and verbs in a sentence.

[0067] For example, to describe the sentence "A cow is standing and chewing the ruminant," we can extract the nouns "cow" and "grass," and the verbs "stand" and "chew."

[0068] A14. Embed synonyms of nouns and synonyms of verbs.

[0069] Generally, word embedding refers to mapping vocabulary to high-dimensional vector space. Noun embedding is used to map nouns to vector space. Verb embedding is used to map verbs to vector space. Sentence embedding is used to convert sentences into fixed-length vectors.

[0070] It should be noted that when embedding, the extracted nouns, verbs, sentences, synonyms of nouns, and synonyms of verbs need to be converted into the form of embedding vectors before the embedding operation is performed.

[0071] Specifically, the embedding vector conversion operation can be performed through a pre-trained Word2vec model (a model for generating word embedding vectors).

[0072] Based on this solution, firstly, it is possible to extract video surface features (global features) from the training videos, as well as motion features and object features (two types of local features) from the training videos. This allows for the extraction of video semantics representing livestock behavioral events from multiple channels. Furthermore, it is possible to obtain descriptive sentences (global features) and verbs and nouns (local features) from the descriptive text, allowing for the extraction of textual semantics representing livestock behavioral events from multiple channels. This means that the semantic information of livestock behavioral events is first enriched from multiple channels, allowing for accurate descriptions of the semantics of the livestock behavioral events to be retrieved from multiple perspectives. Secondly, the descriptive text can also be expanded with synonyms for nouns and verbs, improving the accuracy and robustness of livestock video retrieval based on the input text. Even if the user uses different vocabulary or sentence structures, the correct videos or video frames of livestock containing the target livestock behavioral events will be returned, improving the user experience and the practicality of the retrieval system.

[0073] Specifically, in the embodiment of the present invention, the above-mentioned A14 may specifically include the following steps A41 to A43:

[0074] A41. Aggregate the nouns and their synonyms in the description sentence to generate a noun set corresponding to the first noun search tag.

[0075] The first noun search tag is a predefined noun search tag, such as the actual search tags of the first noun and its synonyms.

[0076] It can be understood that embedding noun synonyms in each sentence containing nouns, that is, merging nouns and noun synonyms and adding them to the set of retrieval tags corresponding to nouns, can enrich the noun semantic information of the descriptive text, such as enriching the description of livestock entities and enriching the location description of livestock behavior events.

[0077] For example, during the search process, by searching for a synonym of noun 1 in the description text as noun 2, and noun 2 being a noun in a preset noun search tag set, noun 2 is determined to be the noun search tag for the search.

[0078] A42. Aggregate the verbs and their synonyms in the description sentence to generate a verb set corresponding to the first verb search tag.

[0079] The first verb search tag is a predefined verb search tag, such as an actual search tag of a synonym of the first verb and the second gerund.

[0080] It can be understood that embedding verb synonyms in each sentence containing a verb, that is, merging the verb and its synonyms and adding them to the synonym set of the retrieval tag corresponding to the verb, can enrich the semantic information of the verb in the sentence, such as enriching the action description of the entity.

[0081] For example, based on the residual attention mechanism, a word (noun or verb) and its synonyms are merged, the weights of the word and its synonyms are calculated through the attention mechanism, the embedding of each synonym is multiplied by the corresponding weight and added to the set of words.

[0082] It can be understood that using the residual attention mechanism to fuse nouns and their synonyms in text sentences, and to fuse verbs and their synonyms in text sentences, can effectively enrich the semantic diversity of nouns and verbs, and enhance the core content semantic information of verbs and nouns in the text, that is, enrich the verb information and noun information describing livestock behavior events.

[0083] A43. Generate a first description sentence retrieval tag based on the first noun retrieval tag, the first verb retrieval tag, and the structure of the extracted description sentence.

[0084] Among them, the first noun retrieval tag, the first verb retrieval tag, and the first description sentence retrieval tag are matched with livestock behavior events in the training video.

[0085] For example, sentence embedding vectors are extracted from description text using a pre-trained Word2vec model. The sentence embedding vector includes the embedding vector of each word in the sentence. The nouns and verbs are parsed from the sentence using the NLTK (Natural Language Toolkit) parser. The nouns parsed by the NLTK parser are converted into noun embedding vectors using the pre-trained Word2vec model. Convert verbs parsed by the NLTK parser into verb embedding vectors Sentence embedding vectors, noun embedding vectors, and verb embedding vectors are obtained through NetVLAD Aggregated into a vector, the embedding vector of the multi-channel features of the extracted text is obtained as follows:

[0086] For example, the sentence embedding vector is represented as represents the set of sentence embedding vectors of the jth sentence in the description text matched by the i-th training video, l se represents the number of words in the jth sentence in the description text matching the i-th training video.

[0087] The noun embedding vector is represented as represents the set of noun embedding vectors of the jth sentence in the description text matched by the i-th training video, l no represents the number of nouns in the jth sentence in the i-th training video.

[0088] The verb embedding vector is represented as represents the set of verb embedding vectors of the jth sentence in the description text matched by the i-th training video, l ve represents the number of verbs in the jth sentence in the i-th training video.

[0089] In an embodiment of the present invention, for the description text matching the animal behavior events in the training video, the description sentences matching the livestock behavior events, the nouns in the description sentences, and the verbs in the description sentences can be first extracted; the extracted description sentences, nouns, and verbs are then converted into description sentence embedding vectors, as well as noun embedding vectors and verb embedding vectors; the description sentence embedding vectors, noun embedding vectors, and verb embedding vectors are respectively aggregated into one vector to obtain an embedding vector description of the extracted text features.

[0090] Specifically, we first extract synonyms of nouns and verbs using the NLTK parser, then use the pre-trained Word2vec model to convert the extracted synonyms of nouns and verbs into embedded forms. We then embed the synonyms of nouns and verbs.

[0091] For example, define the embedding vectors of noun synonyms in each sentence containing a noun as p represents the number of synonyms of the lth noun in the jth sentence of the description text that matches the livestock behavior event in the i-th training video, and no1 represents the first synonym embedding vector of the lth noun.

[0092] For example, define the embedding vectors of verb synonyms in each sentence containing a verb as q represents the number of synonyms of the lth verb in the jth sentence of the description text that matches the livestock behavior event in the i-th training video, and ve1 represents the first synonym embedding vector of the l-th verb.

[0093] In an embodiment of the present invention, a multi-channel feature embedding operation can be performed based on appearance features, object features, motion features, description sentences after content semantic enhancement, a noun set after synonym embedding, and a verb set after synonym embedding.

[0094] In embodiments of the present invention, the multi-channel feature embedding process may be composed of a gated embedding function and a softmax function (activation function). The gated embedding function may recalibrate the strengths of different activations while enhancing nonlinear interactions. For example, the embedding operation may be performed based on the following embedding method.

[0095]

[0096] in, Represents the video features extracted from the training video and the text features extracted from the description text, F∈{app,ob,mo,se,no,ve}. app represents the extracted appearance feature vector, ob represents the extracted object feature vector, mo represents the extracted motion feature vector, se represents the extracted description sentence vector, no represents the extracted noun and noun synonym vector, ve represents the extracted verb and verb synonym vector, are all coefficients learned in advance. Respectively and The dimension of the feature, BatchNorm means batch normalization (Batch Normalization), Represents element-wise multiplication, such as Hadarmard product, σ represents element-wise activation function, such as sigmoid, and φ represents softmax function.

[0097] when When only the text features are represented, The sub-tag i can be ignored. When only the features of the training video are represented, The sub-tag j can be ignored. When representing the features of the training video and the features of the text, The sub-tag i of represents the i-th training video, and the sub-tag j represents the j-th sentence in the description text matching the i-th training video.

[0098] Based on this scheme, by embedding synonyms of verbs and nouns, the semantic information of the descriptive text can be enriched. The richer semantic information describing livestock behavior events, livestock entities and livestock entity movements can improve the accuracy of satisfying the retrieval conditions in the input-based descriptive text retrieval training video and improve the accuracy and robustness of the input-based input text retrieval in livestock videos. Even if the user uses different vocabulary or sentence structure, the correct livestock video or video frame with the target livestock behavior event can be returned, which can improve the user experience and the practicality of the retrieval system.

[0099] Optionally, in the embodiment of the present invention, the temporal semantic enhancement process in S102 may include the following steps B11 and B12:

[0100] B11. Based on the attention mechanism and the extracted surface features, the temporal semantic information describing livestock behavioral events in the training videos is enhanced to obtain the surface features after temporal semantic enhancement.

[0101] For example, the attention mechanism can be used based on formula (1), the extracted appearance features are input into the CNN network to perform time semantic enhancement of the training video, and the parameters of the attention mechanism corresponding to the training video are output. and

[0102]

[0103] in, represents the hth appearance feature of the i-th training video, H represents the number of appearance features of the i-th training video, and W att is the learnable mapping parameter of the attention mechanism, The temporal semantic information of the h-th appearance feature, Represents the weight information of the hth appearance feature corresponding to the i-th training video output by the attention mechanism, represents the temporal semantic information after the hth appearance feature of the i-th training video is enhanced.

[0104] B12. Based on the attention mechanism, GRU (Gated Recurrent Unit) and descriptive sentences, the temporal semantic information of the descriptive sentences of livestock behavior events in the description text is enhanced to obtain the descriptive sentences with enhanced temporal semantics.

[0105] For example, based on formula (2), the temporal semantic information describing the sentence vector can be extracted based on the attention mechanism, GRU and CNN network.

[0106]

[0107] in, represents the average embedding vector of the zth word in the jth sentence of the description text of the i-th training video input to the GRU, W gru represents the learnable mapping parameters of GRU, θ is the other learnable parameters of GRU, Represents the embedded representation information of the jth sentence where the zth verb is output by GRU, The temporal semantic information of the zth verb in the jth sentence of the i-th training video output after temporal semantic enhancement; Represents the weight information of the zth verb in the jth sentence corresponding to the i-th training video output by the attention mechanism, W represents the embedded representation information of the zth verb of the jth sentence corresponding to the i-th training video output by GRU, att is the learnable mapping parameter of the attention mechanism.

[0108] Based on this solution, the temporal information describing livestock behavior in livestock videos can be enhanced using an attention mechanism. Temporal information describing livestock behavior in input text can also be enhanced using the attention mechanism and GRU. This cross-modal temporal semantic enhancement can obtain temporal semantic information about the entire livestock behavior event, enhancing the accuracy of livestock behavior retrieval. For example, a cow normally grazes by lowering its head to gnaw and chew, while a cow ruminates by chewing food stored in its stomach for a long time, without chewing it from the ground or the feed trough. By enhancing the temporal semantics of chewing, the retrieved behavior is determined to be rumination. The chewing behavior duration is determined to be greater than a preset duration, and behavioral events with chewing durations less than the preset duration can be filtered out from the retrieved video, thereby improving the accuracy of cattle rumination retrieval and reducing the probability of retrieving normal cattle grazing behavior as cattle rumination.

[0109] Alternatively, as Figure 2 As shown in , in this embodiment of the present invention, the layered fusion operation in S103 may specifically include the following C11 to C15:

[0110] C11. Fuse the extracted motion features and the extracted object features into motion-object pair features.

[0111] Among them, the fusion of motion features and object features of livestock training videos is the fusion of local features of training videos.

[0112] C12. The extracted motion features, motion-object pair features and temporal semantic enhanced appearance features are fused into video multi-channel features.

[0113] Among them, the fusion of motion features, motion-object pair features and temporal semantic enhanced appearance features is the fusion of global features of the training video.

[0114] C13. Merge the verbs in the verb set and the nouns in the noun set corresponding to the description sentence into word-noun pairs.

[0115] Among them, the word fusion of the verb set and the noun set describing the sentence is the fusion of local text features.

[0116] It should be noted that when fusing verb-noun pairs, the verbs and nouns are usually converted into vector form for fusion to obtain a verb-noun pair in vector form.

[0117] C14. The extracted description sentences, verb-noun pairs, and description sentences enhanced with temporal semantics are integrated into text multi-channel features.

[0118] Among them, the fusion of description sentences, verb-noun pairs and description sentences after time semantic enhancement is the fusion of global text features.

[0119] Similarly, the extracted description sentences and the description sentences after temporal semantic enhancement are converted into vector form and fused with the verb-noun pairs in vector form to obtain the semantically enhanced text multi-channel feature vector form.

[0120] C15. Fusion of video multi-channel features and text multi-channel features as training video-text pair features for livestock behavior event matching.

[0121] In other words, the hierarchical fusion model can first use the attention mechanism to fuse the local-level features of the training video and text to generate new local features, and then use the weighted average method to merge the new local features with the global-level features of the training video and the descriptive text, so as to obtain the essential semantic information of the entire event of livestock behavior and enhance the semantic consistency between different modalities.

[0122] For example, in an embodiment of the present invention, local features and global features of the training video can be fused based on the residual attention mechanism, and local features and global features of the description text can be fused based on the residual attention mechanism.

[0123] Among them, the residual attention mechanism can be described based on the following formula (3).

[0124]

[0125] Among them, u represents the mode of feature fusion, u∈{video,text}, that is, u indicates the training video mode or text mode; h represents the type of feature fusion, h∈{local,global}, that is, h indicates local feature fusion or global feature fusion; y u,h Represents the result of feature fusion output, represents the g-th feature of the feature fusion input, and g represents the total number of features of the feature fusion input.

[0126] Based on this solution, the present invention provides a hierarchical fusion model (HFM). Hierarchical fusion is performed based on the HFM to fuse the multi-channel features of the semantically enhanced training video and the descriptive text, and the semantically enhanced local-level features and global-level features are fused. The local features of the training video and the text can be fused respectively, and the fused local features are fused with their respective global features respectively, so that highly correlated complementary features in each modality are combined to extract global-level (training video appearance and text sentences) features to perceive the entire event of livestock behavior. At the same time, local-level fine-grained features of specific entities and their movements (training video objects and text nouns, training video movements and text verbs) in the event are also paid attention to, and features representing the same attributes are mapped to the same shared space. Comprehensive semantic information of the entire event of livestock behavior can be obtained, which can not only fully highlight the basic semantic information in the training video and the descriptive text, but also reduce redundant information in feature fusion.

[0127] Optionally, in the embodiment of the present invention, the above S104 may include the following D11-D14:

[0128] D11. Determine a first similarity between the video multi-channel feature and the text multi-channel feature.

[0129] D12. Determine the second similarity between the appearance feature after temporal semantic enhancement and the description sentence after temporal semantic enhancement.

[0130] D13. Determine the third similarity of the motion-object pairs and the verb-noun pairs.

[0131] D14. Determine the video-text similarity based on the first similarity, the second similarity, the third similarity, and a predefined similarity function.

[0132] The training video-text similarity function provided by the embodiment of the present invention includes the following formula (5).

[0133] S vt =βS fu +γS app-se +ηS moob-veno Formula (5)

[0134] Among them, S vt Indicates the similarity between video and text, S fu represents the first similarity (i.e., the similarity between the global features of the training video and the global features of the text after hierarchical fusion), S app-se represents the second similarity (i.e., the similarity between the global features of the training video and the global features of the description text before layered fusion), S moob-venorepresents the third similarity (i.e., the similarity between the local features of the training video and the local features of the description text after hierarchical fusion), β represents the similarity fusion coefficient of the first similarity, γ represents the similarity fusion coefficient of the first similarity, and η represents the similarity fusion coefficient of the first similarity, β+γ+η=1.

[0135] It can be understood that the importance of the similarity of features at different angles is different. The importance of the feature similarity of each angle to the overall feature similarity can be learned in advance during the model training process. The importance of the feature similarity of each angle to the overall feature similarity can be adjusted by the weight coefficient corresponding to each feature angle similarity to obtain the overall feature similarity. The weight coefficient corresponding to each feature angle similarity is obtained. The initial values of each similarity fusion coefficient can be the same, and the above-mentioned similarity fusion coefficients can be iteratively optimized based on the retrieval results. Optionally, in an embodiment of the present invention, the similarity coefficient between text and video is calculated to enhance the semantic consistency of different modalities, that is, to enhance the semantic consistency of livestock videos and descriptive texts. Among them, multiple similarities of text and video can be aggregated by weight coefficients.

[0136] Based on this scheme, by training the similarity of the video and the description text from multiple angles, the final training video-text similarity is obtained comprehensively according to the similarity and the importance of the similarity (weight coefficient), which can improve the semantic consistency of different modalities and thus improve the accuracy of cross-modal retrieval.

[0137] Optionally, in the network training process of the text-training video retrieval model of the embodiment of the present invention, in the embodiment of the present invention, the fusion loss function of the training video-text pair is defined as the following formula (6). The fusion loss weight is learned by minimizing the fusion loss function. The fusion loss function uses a bidirectional maximum margin ranking loss to learn the feature embedding vector for training text-video cross-modal retrieval.

[0138]

[0139] Among them, S represents the feature similarity, S∈{S fu ,S app-se ,S moob-veno}, △ is the boundary constant, B represents the number of samples (batch size), δ S (v i ,t b ) represents the similarity score between the feature of the i-th training video and the b-th text label, δ S (v i ,t i ) represents the features and text tags of the i-th training video t i The similarity score, δ S (v b,t i ) represents the training video v b Features and text labels t i The similarity score of , text labels include noun retrieval labels and verb retrieval labels.

[0140] Specifically, positive and negative correlation can indicate whether the text and training video are synchronized. If the livestock training video v i and mark t i Same, or livestock training video v i and text mark t b If both are negative, choose a smaller sample size. Otherwise, choose a larger sample size.

[0141] Optionally, during the network training process, the retrieval loss weights of each feature can be learned by minimizing the multimodal retrieval loss function and the reference multi-task loss fusion based on the multimodal retrieval loss function shown in the following formula (6).

[0142] L=LossFusion(λL fu +μL app-se +σL moob-veno ) Formula (7)

[0143] Among them, L fu represents the first retrieval loss of the video multi-channel feature and the text multi-channel feature, L app-se represents the second retrieval loss of the temporal semantic enhancement appearance feature and the temporal semantic enhancement description sentence, L moob-veno represents the third retrieval loss of the motion-object pair, λ represents the loss weight of the first retrieval loss, μ represents the loss weight of the second retrieval loss, and σ represents the loss weight of the third retrieval loss.

[0144] It can be understood that through iterative operations, when the retrieval accuracy stabilizes, learning is stopped, and the video-text similarity function is determined based on the learned similarity fusion coefficient and hyperparameters. The fusion loss function is determined based on the learned fusion loss weights and hyperparameters. The multimodal retrieval loss function is determined based on the learned retrieval loss weights and hyperparameters. The matching relationship between the video multi-channel features and the text multi-channel features of the livestock behavior event is saved, as well as the retrieval label of the livestock behavior event, the synonym set of the noun retrieval label, the synonym set of the verb retrieval label, and the set of sentence retrieval labels.

[0145] Figure 3 A semantically enhanced logic diagram provided by an embodiment of the present invention, such as Figure 3As shown in the figure, the event is a cow ruminating. The temporal semantic information describing the event in the video is enhanced. The sequence from video frame 1-1 to video frame 1-2 is enhanced to video frame 2-1 and video frame 2-2. Video frame 2-1 precedes video frame 1-1, and video frame 2-2 follows video frame 1-2. The content semantic information of verbs and nouns in the text and the temporal semantic information of sentences are enhanced. Is the cow or the calf ruminating? Synonyms for "cow" include "livestock, native livestock, black ox, deep cattle, yellow cattle, yak, beef cattle, dairy cow..." Synonyms for "calf" include "niu calf, yellow calf, calf, calf, young cow..." Synonyms for "ruminating" include "chewing the cud, chewing back, chewing back grass, foaming back..." "Is the cow or the calf ruminating?" Synonyms for "cow" include "cattle, bull, buffalo, ox, steer," and for "calf" include "little-cow, heifer, bullock." Synonyms for "ruminating" include "chew the cud, regurgitate, remasticate." Temporal semantics are enhanced using the tenses of "is" and "ruminating." A multi-channel feature representation enhancement model (i.e., the aforementioned content semantic enhancement, temporal semantic enhancement, and layered fusion) is used to improve semantic consistency between video text.

[0146] Figure 4This is a schematic diagram of a cross-modal livestock behavior retrieval training model provided in an embodiment of the present invention. Cross-modal livestock behavior retrieval training model 400 includes a data input module 401, a semantic enhancement module 402, a hierarchical feature fusion module 403, and a retrieval training module 404. The semantic enhancement module 402 includes a video content semantic enhancement module 4001, a text content semantic enhancement module 4002, and a temporal semantic enhancement module 4003. The video content semantic enhancement module 4001 includes a video feature extraction module 4010. The text content semantic enhancement module 4002 includes a text feature extraction module 4020 and a feature embedding module 4030. The video feature extraction module 4010 includes an object feature extraction module 4011, a motion feature extraction module 4012, and an appearance feature extraction module 4013. The text feature extraction module 4020 includes a sentence feature extraction module 4021 and a word feature extraction module 4022. The hierarchical feature fusion module 403 includes a video multi-channel feature fusion module 4031, a text multi-channel feature fusion module 4032, and a video-text multi-channel feature fusion module 4033. The retrieval training module 404 includes a similarity determination module 4041 and a loss determination module 4042. The similarity determination module 4041 includes the video-text similarity function in the above formula (5), and the loss determination module 4042 includes the fusion loss function of the training video-text pair in the above formula (6) and the multimodal retrieval loss function in the above formula (7).

[0147] It should be noted that, in an embodiment of the present invention, a cross-modal livestock behavior retrieval training logic diagram is also provided, such as Figure 5 As shown in Figure 5A schematic diagram of a cross-modal livestock behavior retrieval training logic provided by an embodiment of the present invention. For livestock video training, a multi-stream feature representation enhancement model (MFRE) is provided to enhance the content semantic information and temporal semantic information describing livestock behavior events in the video, performing multi-channel feature extraction of video appearance, motion, and objects. For text training, the multi-stream feature representation enhancement model is used to enhance the content semantic information and temporal semantic information describing livestock behavior events in the text, performing multi-channel feature extraction and embedding of text (sentence) nouns, verbs, and synonyms of related nouns and verbs. Specifically, a pre-trained CNN network (i.e., an appearance feature extraction model), an object classification network (i.e., an object feature extraction model), and a pre-trained behavior retrieval network (i.e., an action feature extraction model) were used to extract appearance, motion, and object features from the video. These were then aggregated into three single vector representations corresponding to appearance, motion, and objects using NetVLAD. Simultaneously, a pre-trained Word2vec model and NLTK parser were used to extract features of sentences, nouns, verbs, and their synonyms from the text. These were then aggregated into three single vector representations corresponding to sentences, nouns, and verbs using NetVLAD. The diversity of semantic relevance and a hierarchical fusion model were used to aggregate features across modalities. Specifically, a hierarchical fusion of local and global features from the multimodal video and text was performed, enhancing video-text cross-modal semantics. This enhanced semantic consistency between video and text in the livestock cross-modal behavior retrieval task, reduced redundant information in the fused features, and simultaneously achieved a more comprehensive understanding of the entire event.

[0148] Model application process:

[0149] The following describes an application process of a cross-modal livestock behavior retrieval method provided by an embodiment of the present invention.

[0150] Figure 6 A schematic flow chart of a cross-modal livestock behavior retrieval method provided by an embodiment of the present invention is shown in FIG. Figure 6 As shown in , the following S601 to S606 may be included:

[0151] S601: Acquire a livestock video and input text input by a user.

[0152] The input text includes livestock behavior events.

[0153] For example, the input text may be text inputted based on livestock behavior retrieved according to the needs of a user (eg, a farmer, a veterinarian, or a researcher, etc.).

[0154] It can be understood that due to differences among different users, the descriptions of the input texts input by the users are diverse and different.

[0155] S602: Extract the target sentence, nouns, and verbs of the target sentence from the input text.

[0156] S603: Determine the retrieval tags of nouns and verbs in the target sentence based on the word set corresponding to the livestock behavior retrieval tags.

[0157] The word set corresponding to the livestock behavior search tag includes: a synonym set of the noun search tag and a synonym set of the verb search tag.

[0158] S604: Searching for video features that match the noun retrieval tag, the verb retrieval tag, and the sentence structure of the target sentence from the video-text multi-channel feature set.

[0159] Among them, the video-text multi-channel feature set includes the fusion features of the descriptive text and video corresponding to livestock behavior events after temporal semantic enhancement and content semantic enhancement.

[0160] S605: Retrieve livestock behavior events in the livestock video based on the matched video features.

[0161] Optionally, the search result of the livestock in the livestock behavior event may be a video of a specific livestock having the livestock behavior event, or may be a video frame, which may be set according to user needs and is not specifically limited in the embodiment of the present invention.

[0162] An embodiment of the present invention provides a cross-modal livestock behavior retrieval method. In a first aspect, a livestock video and user-entered input text are obtained. Based on a word set corresponding to a livestock behavior retrieval tag, a retrieval tag for the noun of a target sentence and a retrieval tag for the verb of the target sentence are determined. Then, from a video-text multi-channel feature set, video features that match the noun retrieval tag, the verb retrieval tag, and the sentence structure of the target sentence are searched. Finally, based on the matching video features, livestock behavior events in the livestock video are retrieved. In a first aspect, the word set corresponding to the livestock behavior retrieval tag includes a synonym set for the noun retrieval tag and a synonym set for the verb retrieval tag. Specifically, by using synonyms for nouns and verbs in the input text, noun and verb search tags can be found for retrieval. The noun and verb search tags, along with the sentence structure of the target sentence, accurately describe the essential semantic information of the livestock behavior event to be retrieved. Video features that match the text features are then searched within the video-text multi-channel feature set. Because the video-text multi-channel feature set includes fused features that have been enhanced with temporal and content semantics for both the descriptive text and the video corresponding to the livestock behavior event, the fused features provide a more comprehensive understanding of the entire livestock behavior event and enhance semantic consistency across different modalities. This allows the retrieved video features to more accurately represent the livestock behavior event to be retrieved in the descriptive text. Retrieving livestock videos based on these matched video features effectively improves the accuracy of livestock video-text retrieval, reduces manual control costs, and enables precise and intelligent real-time monitoring and management of large quantities of livestock behavior. For example, based on user-entered text, videos of livestock behavior of interest can be accurately and quickly retrieved from a vast amount of agricultural livestock surveillance videos, locating livestock exhibiting abnormal behavior.

[0163] Optionally, in the livestock behavior cross-modal retrieval method provided in the embodiment of the present invention, the above-mentioned S603 may include the following E1:

[0164] E1. Based on the pre-trained similarity function and the pre-trained retrieval loss function, search for object features, motion features, and appearance features that match the retrieval tag of the noun, the retrieval tag of the verb, and the target sentence structure.

[0165] The motion feature represents the motion of the livestock entity in the livestock behavior event in the video, the object feature represents the livestock entity in the livestock behavior event in the livestock video, and the appearance feature describes the livestock behavior event in the video.

[0166] For example, the video-text similarity between the livestock video and the input text can be calculated based on the pre-trained video-text similarity function of formula (5). The multimodal retrieval loss function can be pre-trained based on formula (7) to perform video-text retrieval and output the retrieval results.

[0167] Based on this scheme, the matching object features, motion features and appearance features can be accurately found through the above-mentioned trained similarity function and pre-trained retrieval loss function.

[0168] Corresponding to the aforementioned method embodiments, the present invention also provides embodiments of a device and a terminal to which the device is applied.

[0169] The embodiment of the cross-modal livestock behavior retrieval method of the present invention can be applied to a computer device, such as a server or a terminal device. The method embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of the cross-modal livestock behavior retrieval device in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 7 As shown in FIG. 1 , a hardware structure diagram of a computer device in which a cross-modal livestock behavior retrieval system according to an embodiment of the present invention resides is shown. Figure 7 In addition to the processor 710, memory 730, network interface 720, and non-volatile memory 740 shown, the server or electronic device where the device 731 is located in the embodiment may also include other hardware according to the actual function of the computer device, which will not be described in detail.

[0170] Figure 8 A schematic diagram of the structure of a cross-modal livestock behavior retrieval device is provided in an embodiment of the present invention. Figure 8As shown in (a) of FIG, the cross-modal livestock behavior retrieval device 800 includes: a data acquisition module 801, a feature extraction module 802, a retrieval tag determination module 803, a feature matching module 804 and a retrieval module 805; the data acquisition module 801 is used to acquire livestock videos and input text input by users, and the input text includes livestock behavior events; the feature extraction module 802 is used to extract the nouns and verbs of the target sentence and the target sentence from the input text; the retrieval tag determination module 803 is used to determine the retrieval tags of the nouns and verbs of the target sentence based on the word set corresponding to the livestock behavior retrieval tags. Retrieval tag; feature matching module 804, used to search for video features that match the noun retrieval tag, verb retrieval tag and sentence structure of the target sentence from the video-text multi-channel feature set; retrieval module 805, used to retrieve livestock behavior events in livestock videos based on the matched video features; wherein the word set corresponding to the livestock behavior retrieval tag includes: a synonym set of the noun retrieval tag, and a synonym set of the verb retrieval tag; the video-text multi-channel feature is a fusion feature after the descriptive text and video corresponding to the livestock behavior event are both enhanced with time semantics and content semantics.

[0171] Optionally, the feature matching module 804 is specifically used to find object features, motion features and appearance features that match the retrieval tags of nouns, the retrieval tags of verbs and the target sentence structure based on a pre-trained similarity function and a pre-trained retrieval loss function; wherein the motion features represent the movement of livestock entities in livestock behavior events in the video, the object features represent livestock entities in livestock behavior events in the livestock video, and the appearance features describe livestock behavior events in the video.

[0172] Optionally, combined Figure 8 (a) in Figure 8As shown in (b) of FIG, the cross-modal livestock behavior retrieval device 800 further includes: a temporal semantic enhancement module 806, a content semantic enhancement module 807, a hierarchical feature fusion module 808, and a retrieval training module 809; the data acquisition module 801 is further used to acquire a training video and a description text matching the livestock behavior events in the training video before acquiring the livestock video and the input text input by the user, wherein the training video is any video in the training video set; the temporal semantic enhancement module 806 is used to perform content semantic enhancement on the training video and the description text, and the content semantic enhancement module 807 is used to perform temporal semantic enhancement on the training video and the description text. The semantic enhancement is performed to obtain the semantically enhanced video features and the semantically enhanced text features; the hierarchical feature fusion module 808 is used to perform hierarchical feature fusion on the semantically enhanced video features and the semantically enhanced text features based on a predefined fusion loss function to obtain the video-text pair multi-channel features after feature fusion; the retrieval training module 809 is used to perform text-video retrieval training based on the predefined text-video similarity function, the predefined multimodal retrieval loss function and the video-text pair multi-channel features, and iteratively train based on the text-video retrieval results to optimize the fusion loss function, the similarity function and the multimodal retrieval loss function.

[0173] Optionally, the content semantic enhancement module 807 is specifically used to: enhance the video local features and video global features describing livestock behavior events in the training video, and enhance the text local features and text global features describing livestock behavior events in the description text.

[0174] Optionally, the content semantic enhancement module 807 is specifically used to: extract object features, motion features and appearance features of the training video; obtain descriptive sentences in the descriptive text, nouns and verbs in the descriptive sentences, extract synonyms of nouns in the descriptive sentences and synonyms of verbs in the descriptive sentences, and embed synonyms of nouns and synonyms of verbs.

[0175] Optionally, the content semantic enhancement module 807 is specifically used to: aggregate nouns and synonyms of nouns in the description sentence to generate a noun set corresponding to the first noun retrieval tag; aggregate verbs and synonyms of verbs in the description sentence to generate a verb set corresponding to the first verb retrieval tag; generate a first description sentence retrieval tag based on the first noun retrieval tag, the first verb retrieval tag and the structure of the extracted description sentence; wherein the first noun retrieval tag, the first verb retrieval tag and the first description sentence retrieval tag match the livestock behavior events in the training video.

[0176] Optionally, the temporal semantic enhancement module 806 is specifically used to: enhance the temporal information describing livestock behavior events in the training video based on the attention mechanism and the extracted surface features, and obtain the surface features after temporal semantic enhancement; enhance the temporal semantic information of the descriptive sentences describing livestock behavior events in the descriptive text based on the attention mechanism, the gated recurrent unit (GRU) and the descriptive sentences, and obtain the descriptive sentences after temporal semantic enhancement.

[0177] Optionally, the hierarchical feature fusion module 808 is specifically used to: fuse the extracted motion features and the extracted object features into motion-object pair features; fuse the extracted motion features, motion-object features and the appearance features after time semantic enhancement into video multi-channel features; based on the structure of the descriptive sentence, fuse the verbs in the verb set and the nouns in the noun set corresponding to the descriptive sentence into verb-noun pairs; fuse the extracted descriptive sentences, verb-noun pairs and the descriptive sentences after time semantic enhancement into text multi-channel features; fuse the video multi-channel features and the text multi-channel features into video-text pair multi-channel features for matching livestock behavior events.

[0178] Optionally, the retrieval training module 809 is also used to: determine a first similarity between video multi-channel features and text multi-channel features; determine a second similarity between the apparent features after temporal semantic enhancement and the description sentences after temporal semantic enhancement; determine a third similarity between motion-object pairs and verb-noun pairs; and determine the video-text similarity based on the first similarity, the second similarity, the third similarity and a predefined similarity function.

[0179] Optionally, the predefined similarity function includes: S vt =βS fu +γS app-se +ηS moob-veno ; Among them, S vt Indicates the similarity between video and text, S fu Indicates the first similarity, S app-se Represents the second similarity, S moob-veno represents the third similarity, β represents the similarity fusion coefficient of the first similarity, γ represents the similarity fusion coefficient of the second similarity, η represents the similarity fusion coefficient of the third similarity, and β+γ+η=1.

[0180] Optionally, the multimodal retrieval loss function includes: L = LossFusion(λL fu +μL app-se +σL moob-veno ); where L fu represents the first retrieval loss of video multi-channel features and text multi-channel features, L app-serepresents the second retrieval loss of the temporal semantic enhancement appearance features and the temporal semantic enhancement description sentences, L moob-veno represents the third retrieval loss of the motion-object pair, λ represents the loss weight of the first retrieval loss, μ represents the loss weight of the second retrieval loss, and σ represents the loss weight of the third retrieval loss.

[0181] Optionally, the fusion loss function includes: Among them, L S represents feature fusion loss, S represents feature similarity, S∈{S fu ,S app-se ,S moob-veno}, △ is the boundary constant, B represents the number of training samples, δ S (v i ,t b ) represents the similarity score between the i-th video and the b-th text label, δ S (v i ,t i ) represents the similarity score between the i-th video and the i-th text label, δ S (v b ,t i ) represents the similarity score between the bth video and the i-th text tag. The text retrieval tags include noun retrieval tags and verb retrieval tags.

[0182] The cross-modal livestock behavior retrieval device provided by the present invention, firstly, enriches the essential semantic information of livestock videos by performing content semantic enhancement and temporal semantic enhancement on livestock videos. It also enriches the essential semantic information of livestock input text by performing content semantic enhancement and temporal semantic enhancement on livestock input text, thereby obtaining essential semantic information about the entire event in different modalities. Secondly, based on the features of the semantically enhanced livestock videos and the features of the semantically enhanced input text, hierarchical fusion is performed to form multi-channel video features and multi-channel text features. This hierarchical fusion combines the complementary features of the two different modalities, video and text, reducing redundant information in the fused features, obtaining a more comprehensive understanding of the entire livestock behavior event, and enhancing semantic consistency across different modalities. This can effectively improve the accuracy of livestock video-text retrieval, reduce the cost of manual control, and achieve precise and intelligent real-time monitoring and management of large amounts of livestock behavior. For example, based on the input text entered by the user, livestock behavior videos of interest can be accurately and quickly retrieved from a large amount of agricultural livestock monitoring videos, and livestock exhibiting abnormal behavior can be located.

[0183] Accordingly, the present invention also provides a cross-modal livestock behavior retrieval device, which includes a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: execute the steps of the cross-modal livestock behavior retrieval method executed by the above-mentioned controller.

[0184] In one embodiment, a computer device is provided, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor implements any step in the above cross-modal livestock behavior retrieval method when executing the computer program.

[0185] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any step in the above cross-modal livestock behavior retrieval method can be implemented.

[0186] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0187] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention. Those of ordinary skill in the art can understand and implement it without paying any creative work.

[0188] The foregoing description describes specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0189] Other embodiments of the present invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow from the general principles of the invention and include common knowledge or customary techniques in the art not claimed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.

[0190] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

[0191] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A cross-modal livestock behavior retrieval method, characterized in that: The method comprises: Acquire a livestock video and input text input by a user, wherein the input text includes livestock behavior events; Extracting a target sentence from the input text, and nouns and verbs of the target sentence; Determining the retrieval tags of nouns and verbs of the target sentence based on the word set corresponding to the livestock behavior retrieval tags; Searching for video features that match the noun retrieval label, the verb retrieval label, and the sentence structure of the target sentence from a video-text multi-channel feature set; Retrieving livestock behavior events in the livestock video based on the matched video features; Among them, the word set corresponding to the livestock behavior retrieval tag includes: a synonym set of the noun retrieval tag, and a synonym set of the verb retrieval tag; the video-text multi-channel feature set includes a fusion feature after the descriptive text and video corresponding to the livestock behavior event are enhanced with time semantics and content semantics; wherein, the time semantics indicates the behavior patterns, habits and relationship between livestock and the environment within a specific time period, and the content semantics indicates the behavior patterns, habits and interactions of livestock.

2. The method according to claim 1, characterized in that The step of searching for video features that match the noun label, the verb label, and the target sentence structure from the video-text multi-channel feature set includes: Based on a pre-trained similarity function and a pre-trained retrieval loss function, searching for object features, motion features, and appearance features that match the retrieval label of the noun, the retrieval label of the verb, and the target sentence structure; The motion feature represents the motion of the livestock entity in the livestock behavior event in the video, the object feature represents the livestock entity in the livestock behavior event in the livestock video, and the appearance feature describes the livestock behavior event in the video.

3. The method according to claim 1 or 2, characterized in that Before obtaining the livestock video and the input text input by the user, the method further includes: Obtaining a training video and a description text matching a livestock behavior event in the training video, wherein the training video is any one video in the training video set; Performing content semantic enhancement on the training video and the description text, and performing time semantic enhancement on the training video and the description text to obtain semantically enhanced video features and semantically enhanced text features; Based on a predefined fusion loss function, performing hierarchical feature fusion on the semantically enhanced video features and the semantically enhanced text features to obtain a multi-channel feature of the video-text pair after feature fusion; Performing text-video retrieval training based on a predefined text-video similarity function, a predefined multimodal retrieval loss function, and the multi-channel features of the video-text pair; Based on iterative training of text-video retrieval results, the fusion loss function, the similarity function and the multimodal retrieval loss function are optimized.

4. The method according to claim 3, characterized in that The content semantic enhancement of the training video and description text includes: The local video features and global video features describing livestock behavior events in the training video are enhanced, and the local text features and global text features describing livestock behavior events in the description text are enhanced.

5. The method according to claim 4, characterized in that The method of enhancing the local video features and global video features describing livestock behavior events in the training video, and enhancing the local text features and global text features describing livestock behavior events in the description text, includes: Extracting object features, motion features, and appearance features of the training video; The description sentences, nouns and verbs in the description sentences are obtained, synonyms of the nouns in the description sentences and synonyms of the verbs in the description sentences are extracted, and the synonyms of the nouns and synonyms of the verbs are embedded.

6. The method according to claim 5, characterized in that The synonyms embedded in the noun and the synonyms embedded in the verb include: Aggregating nouns and synonyms of the nouns in the description sentence to generate a noun set corresponding to the first noun search tag; Aggregating the verbs in the description sentence and synonyms of the verbs to generate a verb set corresponding to the first verb search tag; generating a first description sentence retrieval tag based on the first noun retrieval tag, the first verb retrieval tag, and the extracted structure of the description sentence; The first noun retrieval tag, the first verb retrieval tag, and the first descriptive sentence retrieval tag are matched with livestock behavior events in the training video.

7. The method according to claim 6, characterized in that The performing temporal semantic enhancement on the training video and the description text corresponding to the training video includes: Based on the attention mechanism and the extracted surface features, the temporal information describing the livestock behavior events in the training video is enhanced to obtain surface features with enhanced temporal semantics; Based on the attention mechanism, the gated recurrent unit and the description sentence, the temporal semantic information of the description sentence describing the livestock behavior event in the description text is enhanced to obtain the description sentence with enhanced temporal semantics.

8. The method according to claim 7, characterized in that The step of performing hierarchical feature fusion on the semantically enhanced video features and the semantically enhanced text features to obtain feature-fused video-text pair multi-channel features includes: fusing the extracted motion features and the extracted object features into motion-object pair features; fusing the extracted motion features, the motion-object features, and the temporal semantically enhanced appearance features into a video multi-channel feature; Based on the structure of the description sentence, the verbs in the verb set and the nouns in the noun set corresponding to the description sentence are merged into verb-noun pairs; fusing the extracted description sentences, the verb-noun pairs, and the description sentences enhanced with time semantics into text multi-channel features; The video multi-channel features and text multi-channel features are integrated into video-text pair multi-channel features for livestock behavior event matching.

9. The method according to claim 8, characterized in that The method of performing text-video retrieval training based on a predefined text-video similarity function, a predefined multimodal retrieval loss function, and the multi-channel features of the video-text pair further includes: Determining a first similarity between the video multi-channel feature and the text multi-channel feature; Determining a second similarity between the temporal semantically enhanced appearance feature and the temporal semantically enhanced description sentence; determining a third similarity between the motion-object pair and the verb-noun pair; The video-text similarity is determined based on the first similarity, the second similarity, the third similarity and the predefined similarity function.

10. The method according to claim 9, characterized in that The predefined similarity function includes: S vt =βS fu +γS app-se +ηS moob-veno ; Among them, S vt Indicates the similarity between video and text, S fu represents the first similarity, S app-se Represents the second similarity, S moob-veno represents the third similarity, β represents the similarity fusion coefficient of the first similarity, γ represents the similarity fusion coefficient of the second similarity, η represents the similarity fusion coefficient of the third similarity, and β+γ+η=1.

11. The method according to claim 10, characterized in that The multimodal retrieval loss function includes: L = LossFusion (λL fu +μL app-se +σL moob-veno ); Among them, L fu represents the first retrieval loss of the video multi-channel feature and the text multi-channel feature, L app-se represents the second retrieval loss of the temporal semantic enhancement appearance feature and the temporal semantic enhancement description sentence, L moob -veno represents the third retrieval loss of the motion-object pair, λ represents the loss weight of the first retrieval loss, μ represents the loss weight of the second retrieval loss, and σ represents the loss weight of the third retrieval loss.

12. The method according to claim 11, characterized in that The fusion loss function includes: Among them, L S represents feature fusion loss, S represents feature similarity, S∈{S fu ,S app-se ,S moob-veno }, Δ is the boundary constant, B represents the number of training samples, δ S (v i ,t b ) represents the similarity score between the i-th video and the b-th text label, δ S (v i ,t i ) represents the similarity score between the i-th video and the i-th text label, δ S (v b ,t i ) represents the similarity score between the bth video and the i-th text tag, where the text retrieval tags include noun retrieval tags and verb retrieval tags.

13. A cross-modal livestock behavior retrieval device, characterized in that: The cross-modal livestock behavior retrieval device includes: a data acquisition module, a feature extraction module, a retrieval tag determination module, a feature matching module and a retrieval module; The data acquisition module is used to acquire livestock videos and input text input by the user, wherein the input text includes target livestock behavior events; The feature extraction module is used to extract the target sentence, the nouns of the target sentence and the verbs of the target sentence from the input text; The retrieval tag determination module is used to determine the retrieval tags of nouns and verbs of the target sentence based on the word set corresponding to the livestock behavior retrieval tags; The feature matching module is used to search for video features that match the noun retrieval tag, the verb retrieval tag, and the sentence structure of the target sentence from the video-text multi-channel feature set; The retrieval module is used to retrieve livestock behavior events in the livestock video based on the matched video features; wherein the word set corresponding to the livestock behavior retrieval tag includes: a synonym set of the retrieval tag of the noun, and a synonym set of the retrieval tag of the verb; the video-text multi-channel feature is a fusion feature after the descriptive text and video corresponding to the livestock behavior event are enhanced with time semantics and content semantics; wherein the time semantics indicates the behavior patterns, habits and relationship between the livestock and the environment within a specific time period, and the content semantics indicates the behavior patterns, habits and interactions of the livestock.

Citation Information

Patent Citations

  • Livestock climbing behavior labeling method and device, electronic equipment and storage medium

    CN115661717A

  • Video description method based on cross-modal retrieval semantic enhancement and storage medium

    CN118153561A