Multi-modal large model-based phenotypic feature detection and identification method for low-yield chickens

By combining a large multimodal model with visual and textual information, the nonlinear changes of oligotrophic chickens are analyzed, solving the problem of low accuracy in identifying oligotrophic chickens in existing technologies. This enables efficient and high-precision identification of oligotrophic chickens, and promotes the application of inspection robots in large-scale chicken houses.

CN120808405AActive Publication Date: 2025-10-17CHINA AGRI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511316724.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-17
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing technologies have low recognition accuracy when detecting oligotrophic chickens, are easily affected by environmental factors, and have difficulty analyzing the nonlinear changes between the phenotypic characteristics of chickens and their age. This results in low recognition efficiency and high labor intensity, limiting the application of inspection robots in large-scale chicken houses.

Method used

A phenotypic feature detection method for oligotrophic chickens based on a multimodal large model is adopted. Combining visual and text information, age information is introduced to analyze the nonlinear changes of oligotrophic chickens. Through multiple rounds of instruction-following data training model, efficient and high-precision oligotrophic chicken identification is achieved.

Benefits of technology

It has achieved efficient and high-precision identification of oligotrophic chickens, promoted the widespread application of inspection robots in large-scale chicken houses, and provided key technical support for intelligent egg-laying poultry farming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808405A_ABST
    Figure CN120808405A_ABST
Patent Text Reader

Abstract

The invention, which relates to the technical field of computer vision and artificial intelligence, discloses a method for detecting and identifying phenotypic characteristics of low-yield chickens based on a multi-modal large model, comprising the following steps: acquiring image data and day age of to-be-identified cage-rearing chickens in a chicken house; and taking the related problem text of the recognition of the low-yield chickens, the image data of the to-be-recognized cage-rearing chickens and the day ages of the to-be-recognized cage-rearing chickens as input of a pre-trained accurate recognition multi-modal large model for the low-yield chickens, and obtaining recognition results and coordinates of the low-yield chickens. According to the method, high-efficiency and high-precision identification of the low-laying chickens in the large-scale henhouse is realized, a theoretical basis is provided for realizing the accurate identification function of the low-laying chickens of the inspection robot for the large-scale henhouse, the wide application of the inspection robot in the large-scale henhouse is promoted, and a key technical support is provided for intelligent breeding of laying fowls.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision and artificial intelligence, and in particular to a kind of oligogenic chicken phenotype feature detection and identification method based on multi-modal large model. BACKGROUND

[0002] In the process of large-scale poultry breeding, the detection and selection of oligogenic chickens (which have a much lower egg production rate than normal during the egg-laying period) cannot be ignored. According to statistics, oligogenic chickens account for about 3% of large-scale chicken flocks, severely affecting breeding efficiency and industry economic benefits. According to data such as the number of laying hens in 2024, average egg prices, average feed consumption, water consumption, and carbon emissions, "eating only feed and not laying eggs" caused by oligogenic chickens consumes about 230,000 tons of feed and 214 million tons of water per year, emits 380,000 tons of CO2, and directly causes economic losses of nearly 10 billion yuan.

[0003] Currently, more than 80% of laying hens in China are raised in intensive cage systems, and some large-scale farms use the experience method of "one touch and five observations", especially "night touch", to identify oligogenic chickens in the flock and eliminate them in a timely manner to improve breeding efficiency. However, due to the high stress on the chickens, low identification efficiency, and high labor intensity, this method cannot be widely applied. Therefore, it is necessary to detect oligogenic chickens through intelligent means to reduce economic losses.

[0004] Some existing intelligent detection methods for oligogenic chickens rely solely on a single phenotype (such as comb color and size) for judgment. Such detection methods face two major limitations in practical application scenarios: first, detection models based on a single phenotype feature are highly sensitive and easily affected by environmental factors; second, in intensive cage systems, chickens often block each other, and once the key detection features are blocked, the model is prone to misjudgment. In addition, existing detection technologies mainly analyze the phenotype of the chicken, but have not effectively analyzed the non-linear relationship between the chicken's phenotype features and age, and still rely on manual experience to determine the best identification time, which to some extent limits the development of end-to-end models and restricts the application of oligogenic chicken identification functions in actual chicken coops. SUMMARY

[0005] To address the above deficiencies in the prior art, the present application provides a kind of oligogenic chicken phenotype feature detection and identification method based on multi-modal large model, which solves the problem of low recognition accuracy in the prior art.

[0006] To achieve the above-mentioned purposes, the technical solution adopted by the present application is as follows: a kind of oligogenic chicken phenotype feature detection and identification method based on multi-modal large model, comprising: obtaining image data and age of the cage chickens to be identified in the chicken coop; The relevant question text for the identification of poor-born chickens, the image data of caged chickens to be identified, and their age are used as the input of the pre-trained multimodal large model for accurate identification of poor-born chickens to obtain the identification results and coordinates of poor-born chickens.

[0007] The beneficial effects of the present invention are: Based on the multimodal large model, the visual-text joint analysis of the phenotypic characteristics of oligotrophic chickens is carried out, and text information is used to enhance the visual feature signals of oligotrophic chickens, effectively avoiding the problems of fuzzy recognition features and inaccurate analysis of single modality; age information is introduced to analyze the nonlinear changes of phenotypic characteristics of oligotrophic chickens at different physiological stages, and the egg production rate auxiliary signal is introduced during reasoning to enhance the model's recognition sensitivity, ultimately achieving efficient and high-precision identification of oligotrophic chickens in large-scale chicken houses, providing a theoretical basis for the accurate identification function of oligotrophic chickens by large-scale chicken house inspection robots, helping to promote the widespread application of inspection robots in large-scale chicken houses, and providing key technical support for the intelligent breeding of laying poultry. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A flow chart of a method for detecting and identifying phenotypic characteristics of oligotrophic chickens based on a multimodal large model is provided in an embodiment; Figure 2 A schematic diagram of the multimodal large model for accurately identifying oligotrophic chickens provided in the embodiment; Figure 3 Schematic diagram of the marking and segmentation of the phenotypic parts of oligotrophic chickens provided in the embodiment; Figure 4 An attention heat map for discriminating oligotrophic chickens based on phenotypic characteristics is provided in the embodiment. DETAILED DESCRIPTION

[0009] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0010] like Figure 1 As shown, in one embodiment of the present invention, a method for detecting and identifying the phenotypic characteristics of oligotrophic chickens based on a multimodal large model includes the following steps: S1. Build a multi-modal model for accurate identification of oligotrophic chickens: like Figure 2As shown, the oligoparent chicken accurate identification multi-modal large model includes a visual encoder (the visual encoder is a double-branch structure including a VGGT and a CLIP image encoder), a projection layer, and a base large language model (the base large language model can be a GPT-4, Llama, Qwen 2.5, etc. mainstream architecture, so that the oligoparent chicken multi-modal large model can be upgraded and iterated synchronously with the upgrade of the base model, maintaining its high performance). The day-old information projection layer, the depth projection layer, and the visual text projection layer are all projection layers, wherein the day-old information projection layer is a linear layer structure, and the depth projection layer and the visual text projection layer are both two-layer MLP structures.

[0011] S2, acquiring image data of cage chickens in a chicken house, generating description image text pair data based on a general large model, and converting the image text pair into a high-quality instruction following data set: Acquiring image data of cage chickens in a chicken house; generating description image text pair data based on a general large model according to the acquired image data, and converting the image text pair data into a high-quality instruction following data set. In constructing the high-quality instruction following data set, initial description image text pair data is generated based on ChatGPT and Qwen 2.5 VL, specifically including: using image titles, bounding boxes, day-old information of oligoparent chickens, and coordinates of key phenotype parts (such as Figure 3 、 Figure 4 As shown, the comb, beak, eye, body area, and chicken claw can be used as the phenotype part for identifying oligoparent chickens) as symbolic representation, and the image is encoded into a sequence recognizable by the large language model. Three types of instruction following data, dialogue, detailed description, and complex reasoning, are designed, and the description is fine-tuned by experienced veterinarians according to the format to construct a high-quality image-text pair data set, which requires clear description of the characteristics of different phenotype parts of each chicken and the ability to make a judgment on whether it is an oligoparent chicken based on these phenotype characteristics. Query the general large model through context learning to generate high-quality instruction following data.

[0012] Context type 1 (image title): There are eight chickens in a cage, they are full of energy, lively in appearance, and look healthy.

[0013] There is an egg-laying bar in front of the chicken coop, with 6 eggs on it.

[0014] In a chicken coop, a group of chickens are blocking each other, among which three are drinking water.

[0015] There is a chicken lying down in the chicken group, looking bloated in the abdomen and having no eyes, suspected to be an oligoparent chicken.

[0016] The chickens in the chicken coop are competing for feed, and they look like some of their feathers have fallen off.

[0017] Context type 2 (supervision signal coordinates): Day of age {440 days}, target frame for low-producing chickens [0.681, 0.242, 0.774, 0.694], top of comb for low-producing chickens [0.758, 0.413, 0.845, 0.69], eyes for low-producing chickens [0.63, 0.222, 0.686, 0.516], beak tip for low-producing chickens [0.755, 0.413, 0.846, 0.692], muffs for low-producing chickens [0.384, 0.696, 0.485, 0.914], shanks for low-producing chickens [0.1, 0.497, 0.173, 0.579], laying cage [0.261, 0.101, 0.787, 0.626].

[0018] The generated instruction following dataset includes three types: dialogue, detailed description, and complex reasoning, for example: Type 1: Dialogue Question: Please describe the scene in the picture. Answer: The picture shows a chicken coop with limited interior space, slightly dim light, and several chickens moving inside, which may be in a productive chicken house.

[0019] Question: How many chickens are in the picture? Answer: There are eight or nine chickens in the picture, four of which are near the feeding trough, two are active in the middle area, and two or three are in the dark corner, one of which is severely obscured.

[0020] Type 2: Detailed description There are eight brown egg-laying hens in the wire chicken coop. Three chickens on the left side are fighting for food in the feeding trough, with bright red combs standing upright; two chickens in the middle have soft abdomens and disheveled feathers standing upright; a chicken in the corner on the right side is curled up: the comb is pale and atrophied, the beak and claws are still yellow, the feathers are close to the body, and the head is low and the neck is curled.

[0021] There are eight hens in the chicken coop. Two chickens in the front are pecking, with bright red combs; four chickens in the middle are active, with full abdomens; one chicken in the back is standing against the wall with drooping wings, with a bright yellow beak, a white comb, and disheveled feathers.

[0022] Type 3: Complex reasoning Question: Please tell me if there are low-producing chickens in this image? If so, output their coordinates and the basis for your judgment.

[0023] Answer: Seven roosters in the flock have red combs, white legs and feet, and wide abdomens; one rooster is crouched in the corner of the cage, with a small and pale comb, bright yellow legs and feet, and a narrow and hard abdomen, with no feather loss. The rooster in the corner has a pale and atrophic comb, a yellowish face, and still yellowish legs and feet, with smooth and shiny feathers, which is likely to be a low-producing hen, with coordinates [0.660, 0.420, 0.486, 0.615]. Since the current age is close to 400 days, the nutrients ingested by the low-producing hen after not laying eggs are mainly used for maintaining its own metabolism, so it shows the phenomenon of smooth and shiny feathers. At the same time, the yellow pigment ingested in the feed cannot be transferred to the egg yolk, and thus is deposited in the beak, legs and feet, etc., showing a distinct bright yellow color.

[0024] S3, training a low-producing hen precise identification multi-modal large model according to instruction following data sets; The training of the low-producing hen precise identification multi-modal large model includes a first stage and a second stage; in the first stage, the visual encoder and the base large language model are frozen, and the projection layer is trained using single-round instruction following data (such as "please describe the image" + artificial annotation description) to make the visual features and text features expressed in the same semantic space; In the second stage, the visual encoder is frozen, and the projection layer and the base large language model are trained using the generated multi-round instruction following data (including description, reasoning, dialogue, etc.) in a self-recurrent manner, while updating the weights of the projection layer and the base large language model; The specific training process of the second stage is: The generated high-quality instruction data set is used as the input of the low-producing hen precise identification multi-modal large model, and its expression is:

[0025] wherein, is the input sequence, is the instruction text, is the target answer; The hidden state is generated :

[0026] wherein, represents the processing process of the structure; The token prediction is performed according to the hidden state, and its expression is:

[0027] wherein, represents the probability of predicting the next token under the given condition , represents the token of the t th step, represents the time stept all tokens before the current token; denotes a softmax function, denotes the output layer weight of the base large language model; The token prediction probability maximizes the likelihood probability of the target answer sequence, and its expression is:

[0028] wherein, denotes the likelihood probability of the target answer sequence, denotes the total number of tokens of the target answer, denotes a natural logarithm, denotes all parameters of the projection layer and the base large language model, denotes the i token of the target answer, i denotes the token index.

[0029] S4, taking the related problem text of the identified few-production chickens, the image data of the to-be-identified cage chickens and the day-old of the to-be-identified cage chickens as the input of the pre-trained few-production chicken accurate identification multi-modal large model, obtaining the identification result and coordinates of the few-production chicken; Specifically: a. Collecting image data of to-be-identified cage chickens in the chicken house by a patrol robot, inputting the image data of to-be-identified cage chickens into a visual encoder; b. The VGGT is used to predict the depth information of the image data, and the output feature map of the DPT head in the VGGT is extracted as the depth feature , and its expression is:

[0030] wherein, denotes the processing process of the image data by the VGGT, denotes a feature map with a height of , a width of , and a channel number of ; and c. The CLIP image encoder is used to obtain the RGB image feature of the image data, and its expression is:

[0031] wherein, denotes the processing process of the image data by the CLIP image encoder; denotes the number of image data patches, denotes the original dimension of the CLIP feature, denotes a feature map with a row number of , and a column number of​ matrix feature map; d. mapping the depth feature to the visual feature space through the depth projection layer to obtain a mapped feature , whose expression is:

[0032] wherein, , respectively represent the weight matrix of the first layer MLP structure and the second layer MLP structure in the depth projection layer; represents an activation function; , respectively represent the bias vector of the first layer MLP structure and the second layer MLP structure in the depth projection layer; represents the token number of the depth feature; represents a feature matrix with rows and columns; e. aligning and channel splicing the mapped feature with the RGB image feature to obtain a visual feature containing depth information, whose expression is:

[0033] wherein, represents feature fusion; represents a feature matrix with rows and columns; f. discretely grouping the physiological stages of the laying hens to obtain day age grouping of 5 different stages;

[0034] wherein, is a grouping index, is the minimum day age value in the data set, is the day age grouping interval, represents day age metadata, represents discretely grouping the data set .

[0035] The day age grouping can allow partial overlap within each grouping to alleviate the influence of abnormal distribution of data on model overfitting during training, for example: Discretely grouped: {AGE_BINS=[(16,19),(20,28),(29,43),(44,55),(56,100)]}; Grouping with overlap: {AGE_BINS_OVERLAP=[(16,19),(18,28),(27,43),(42,55),(54,100)]}.

[0036] g. The token embedding of different stages of age groups is obtained by projecting the age information layer, and the age feature aligned with the visual feature is obtained. ; wherein, represents an embedding function, represents age metadata, represents the obtained different stages of age groups; represents a feature matrix with a row number of 1 and a column number of ; h. The age feature is concatenated with the visual feature to obtain the visual feature with age information concatenated, so that the projected visual feature can be “understood” by the language model, and the expression is:

[0037] wherein, represents a feature matrix with a row number of and a column number of ; i. The text description data of the corresponding question is token embedded to obtain the text feature ; j. The visual feature with age information concatenated is input into the visual-text projection layer and the text feature to obtain the feature , and the expression is: ; wherein, , respectively represent the weight matrix of the first layer MLP structure and the second layer MLP structure in the visual-text projection layer; , respectively represent the bias vector of the first layer MLP structure and the second layer MLP structure in the visual-text projection layer; is the unified feature dimension of the oligopotent chicken accurate identification multi-modal large model; represents a feature matrix with a row number of and a column number of ; k. The feature is concatenated with the text feature The channel splicing is performed to obtain cross-modal sequence information , and the expression is:

[0038] , wherein, represents the length of the text sequence, represents a feature matrix with a row number of and a column number of .

[0039] l. The aligned cross-modal sequence information is taken as an input of the base large language model to obtain the identification result and coordinates of the oligopregnant hen.

[0040] The input of the oligopregnant hen precise identification multi-modal large model is the image data of the cage hen to be identified, the day-old information, and the corresponding problem description. For example, the visual features, day-old metadata, and text features (such as the question "describe this picture") are spliced to form a joint input sequence: {[IMG_1, IMG_2,..., IMG_256] + [age]+ [Q1, Q2,...,Qn]}.

[0041] When the oligopregnant hen precise identification multi-modal large model identifies the oligopregnant hen, a sensitivity coefficient of the egg-laying rate of each cage is introduced to dynamically adjust the identification threshold of the oligopregnant hen in different cages. When the single cage egg-laying rate is lower than the average egg-laying rate, the probability of the existence of the oligopregnant hen is increased, and the sensitivity coefficient of the corresponding cage is increased, otherwise the sensitivity coefficient of the corresponding cage is decreased.

[0042] The identification threshold of the oligopregnant hen is adaptively adjusted by introducing a temperature parameter to scale the original logits or dynamic text prompts. The logits probability distribution scaled by the temperature coefficient T is:

[0043]

[0044]

[0045]

[0046] , wherein, is the logits probability distribution scaled by the temperature coefficient T, is a natural constant, represents the logits scaled by the temperature coefficient T, represents the number of logits, represents the original logits, represents the temperature parameter, represents the sensitivity adjustment coefficient, represents the deviation of the egg laying rate, represents the current cage egg laying rate, represents the average egg laying rate.

[0047] Dynamic text prompts, for example: sensitivity prompts: {Single cage egg laying rate ≤ 50%: sensitivity = "Identify any possible low-producing chicken, even if not very sure"} {50% ≤ single cage egg laying rate ≤ 70%: sensitivity = "Identify individuals who may appear to be low-producing chickens"} {Single cage egg laying rate ≤ average egg laying rate: sensitivity = "Identify obvious low-producing chickens"} {Others: sensitivity = "Only identify very certain low-producing chickens"} Prompt prompt word: "[AGE_{age}] The egg laying rate of this cage is lower than the average value {}%, please {sensitivity}".

[0048] The multimodal large model involved in the present application can be deployed in an inspection robot device, and through setting an inspection task, efficient inspection of low-producing chickens can be realized throughout the day. At the same time, the model has strong expansibility, and in the future, only by adding the image and text data of dead chickens and weak chickens can the dead chicken and weak chicken recognition function of the model be expanded. Common abnormal chicken data is unified into the multimodal large model framework, and the vertical multimodal large model specific to poultry breeding is fine-tuned.

[0049] In summary, the present application realizes efficient and high-precision identification of low-producing chickens in large-scale chicken coops, provides a theoretical basis for realizing the precise identification function of low-producing chickens for large-scale chicken coop inspection robots, helps to promote the wide application of inspection robots in large-scale chicken coops, and provides key technical support for egg poultry intelligent breeding.

Claims

1. A method for detecting and identifying phenotypic characteristics of oligotrophic chickens based on a multimodal large model, characterized in that: include: Obtain image data and age of caged chickens to be identified in the chicken house; The relevant question text for the identification of poor-born chickens, the image data of caged chickens to be identified, and their age are used as the input of the pre-trained multimodal large model for accurate identification of poor-born chickens to obtain the identification results and coordinates of poor-born chickens.

2. The method according to claim 1, characterized in that The multimodal large model for accurate identification of widowed chickens includes a visual encoder, a projection layer, and a base large language model connected in sequence.

3. The method according to claim 2, characterized in that The visual encoder has a dual-branch structure, one branch is VGGT and the other branch is CLIP image encoder; the input of the visual encoder is used as the input of VGGT and CLIP image encoder, and the output of VGGT and CLIP image encoder is used as the output of the visual encoder.

4. The method according to claim 3, characterized in that The age information projection layer, depth projection layer and visual text projection layer are all projection layers, among which the age information projection layer is a one-layer MLP structure, and the depth projection layer and visual text projection layer are both two-layer MLP structures.

5. The method according to claim 4, characterized in that The specific method for obtaining the identification results and coordinates of the widowed chicken is: using the projection layer to gradually align the image and text feature spaces to obtain cross-modal sequence information; using the aligned cross-modal sequence information as the input of the base large language model to obtain the identification results and coordinates of the widowed chicken.

6. The method according to claim 5, characterized in that The specific method of using the projection layer to gradually align the image and text feature spaces is as follows: The image data of caged chickens to be identified Input visual encoder; The depth information of the image data is predicted by VGGT, and the output feature map of the DPT head in VGGT is extracted as the depth feature , whose expression is: in, Represents the VGGT processing process of image data, Indicates the height , width is , the number of channels is Feature map of Obtain the RGB image features of the image data through the CLIP image encoder , whose expression is: in, Indicates the processing of image data by the CLIP image encoder; Indicates the number of image data patches, represents the original dimension of CLIP features, Indicates the number of rows , the number of columns is Matrix feature map of ; The depth feature is mapped to the visual feature space through the depth projection layer to obtain the mapping feature , whose expression is: in, 、 Represent the weight matrices of the first and second MLP structures in the depth projection layer respectively; express Activation function; 、 Represent the bias vectors of the first and second MLP structures in the depth projection layer respectively; The number of tokens representing deep features; Indicates the number of rows , the number of columns is The characteristic matrix of Mapping features and RGB image features Align and stitch channels to obtain visual features containing depth information , whose expression is: in, Indicates feature fusion; Indicates the number of rows , the number of columns is The characteristic matrix of The physiological stages of laying hens are discretized and grouped to obtain age groups at different stages; The age information projection layer is used to group the age groups at different stages and embed tokens to obtain the visual features. Aligned age characteristics , whose expression is: ; in, represents the embedding function, Indicates age metadata. Indicates the age groups obtained at different stages; Indicates that the number of rows is 1 and the number of columns is The characteristic matrix of Age characteristics and visual features Perform channel splicing to obtain visual features with age information , whose expression is: in, Indicates the number of rows , the number of columns is The characteristic matrix of Token embedding of text description data to obtain text features ; The visual features with age information are spliced ​​together Input visual text projection layer and text features are dimensionally aligned to obtain features , whose expression is: ; in, 、 Represent the weight matrices of the first and second MLP structures in the visual text projection layer respectively; 、 Respectively represent the bias vectors of the first and second MLP structures in the visual text projection layer; Unified feature dimensions for multimodal large models to accurately identify oligotrophic chickens; Indicates the number of rows , the number of columns is The characteristic matrix of The features With text features Perform channel splicing to obtain cross-modal sequence information , whose expression is: in, Indicates the length of the text sequence, Indicates the number of rows , the number of columns is The feature matrix of .

7. The method according to claim 6, characterized in that The specific training process of the multimodal large model for accurately identifying oligotrophic chickens is as follows: Acquire image data of caged chickens in a chicken house; use a universal large-scale model to generate descriptive image-text pairs based on the acquired image data, and convert the image-text pairs into a command-following dataset; input the command-following dataset into a multimodal large-scale model for accurate identification of oligotrophic chickens for training; The method of using the universal large model to generate descriptive image text pairs based on the acquired image data specifically includes: using the title, bounding box, age of the oligofatty chickens, and coordinates of key phenotypic parts of the image data as symbolic representations, encoding the acquired image data into a sequence recognizable by the universal large language model; querying the universal large model through context learning based on the encoded sequence to generate a command following dataset, which includes three types: dialogue, detailed description, and complex reasoning; and having experienced breeders adjust the generated command following set to obtain the final command following dataset.

8. The method according to claim 7, characterized in that The training of the large multimodal model for accurate identification of rare chickens consists of the first and second phases. The first phase freezes the visual encoder and the base large language model, uses a single round of instruction-following data training, and trains only the projection layer, so that visual features and text features are expressed in the same semantic space. In the second stage, the visual encoder is frozen and the projection layer and the base language model are trained autoregressively using the generated multi-round instruction following data, while the weights of the projection layer and the base language model are updated. The specific training process of the second stage is: The generated instruction dataset is used as the input of the multimodal large model for accurate identification of oligotrophic chickens, and its expression is: in, is the input sequence, is the instruction text, For the target answer; Generate hidden state : in, express Structural processing; Token prediction is performed based on the hidden state, and its expression is: in, Indicates that under given conditions Predict the probability of the next token, Represents the return t Step token, Represents the time step t All tokens before; represents the softmax function, Represents the output layer weight of the base large language model; Maximize the likelihood probability of the target answer sequence based on the token prediction probability, and its expression is: in, represents the likelihood probability of the target answer sequence, The total number of tokens representing the target answer, represents the natural logarithm, Represents all parameters of the projection layer and the base language model, The target answer i tokens, i Indicates the token index.

9. The method according to claim 8, characterized in that When using a large multimodal model for accurately identifying poorly laid chickens, the egg production rate of each cage is introduced to adjust the sensitivity coefficient, thereby dynamically adjusting the judgment threshold for poorly laid chickens in different cages; when the egg production rate of a single cage is lower than the average egg production rate, the probability of poorly laid chickens increases, and the sensitivity coefficient of the corresponding cage is increased; otherwise, the sensitivity coefficient of the corresponding cage is reduced.

10. The method according to claim 9, characterized in that Adaptively adjust the threshold for determining oligotrophic chickens by introducing temperature parameters to scale the original logits or dynamic text prompts; The logits probability distribution after scaling by the temperature coefficient T is: in, is the logits probability distribution after scaling by the temperature coefficient T, is a natural constant, represents logits scaled by the temperature coefficient T, represents the number of logits, represents the original logits, represents the temperature parameter, represents the sensitivity adjustment coefficient, Indicates the egg production rate deviation, Indicates the current egg production rate of the chicken cage. Indicates the average egg production rate.

Citation Information

Patent Citations

  • Integrated equipment for detecting weak and dead cage-rearing laying hens

    CN117671586A

  • Disease and pest knowledge generation type question answering method based on big and small model collaboration

    CN118551845A

  • Free-range chicken multi-scale target detection method and device based on large model

    CN118675199A

  • Chicken flock state inspection monitoring system and method

    CN119989281A

  • Abnormality determination system, abnormality determination device and abnormality determination method

    JP2017192316A