Pet-searching geographic information matching method based on machine learning and natural language processing

By applying machine learning and natural language processing technology on the pet hunting platform, the visual and semantic features of pets are extracted, and dynamic matching degree is calculated based on space-time information to generate thermal maps, the problems of inaccurate matching of existing platforms and slow information updates are solved, and the efficiency and success rate of pet hunting are improved.

CN120180157AActive Publication Date: 2025-06-20SICHUAN NORMAL UNIV

Patent Information

Application Number
CN202510667746.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing pet search platform is difficult to accurately match lost pet information, and the information update speed is slow, which makes users spend a lot of time and energy on finding matching information in massive information, and the success rate is not high.

Method used

Using a method based on machine learning and natural language processing, the visual features of pet images and semantic features of text description are extracted through deep learning models, the lost spatiotemporal information is encoded into spatiotemporal joint vectors, and these features are fused to calculate dynamic matching degrees, a visual thermal map is generated, and an interactive map interface is output.

Benefits of technology

It improves the efficiency and accuracy of pet search, enhances the reliability of predicting the range of activities of lost pets, reduces users' time and energy investment in information retrieval, and improves the success rate of pet search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180157A_ABST
    Figure CN120180157A_ABST
Patent Text Reader

Abstract

The invention relates to a pet-searching geographic information matching method based on machine learning and natural language processing. The method comprises the following steps: acquiring a pet image, text description and loss space-time information input by a user, wherein the loss space-time information comprises loss position coordinates and loss time; extracting visual features of the pet image and semantic features of text description through a deep learning model; coding the lost spatio-temporal information into a spatio-temporal joint vector, fusing the visual features and the semantic features, and calculating a dynamic matching degree; and generating a visual thermodynamic map according to the dynamic matching degree, performing screening according to a preset screening threshold, and outputting an interactive map interface. According to the method, visual, semantic and spatio-temporal information is comprehensively analyzed through machine learning and natural language technologies, the pet searching accuracy and efficiency are remarkably improved, the problems of insufficient information utilization and low matching accuracy in a traditional pet searching method are solved, and powerful support is provided for quickly finding a lost pet.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to a lost pet geographical information matching method based on machine learning and natural language processing. Background Art

[0002] With the rapid development of technologies in the field of artificial intelligence, in the field of pet searching, there has emerged a pet searching platform technology based on information technology. This technology can attempt to match relevant records of lost pets and found pets by integrating information provided by users. Compared with the traditional offline pet searching method, it breaks through the limitations of time and space and can quickly spread pet searching information over a larger range. However, the existing pet searching platforms can only provide a platform for users to upload pet searching notices. On the one hand, the pet owner usually can only search for relevant information by keywords such as the general lost location. When the description entered by the user does not exactly match the description of the publisher, it is very difficult to accurately match the relevant missing pet information. On the other hand, the lost pet information uploaded by users is not processed in a timely manner, resulting in a waste of a large amount of valid information, and the information update speed is slow, making many of the information retrieved by users may have become outdated. For example, the pet may have been found but the information still shows as lost on the platform, resulting in users having to search for information related to their pets manually in a vast amount of information, which not only consumes a lot of time and energy but also is prone to omissions and errors, leading to a low success rate of pet searching. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a lost pet geographical information matching method based on machine learning and natural language processing to improve the efficiency and accuracy of pet searching and enhance the reliability of predicting the activity range of lost pets.

[0004] In a first aspect, the present application provides a lost pet geographical information matching method based on machine learning and natural language processing, including: Obtaining a pet image, a text description, and lost space-time information input by a user, where the lost space-time information includes lost position coordinates and lost time; Extracting visual features of the pet image and semantic features of the text description through a deep learning model, where the visual features include breed, color, body shape, and special marks, and the semantic features include breed entities, color modifiers, and behavioral features; Encoding the lost space-time information into a space-time joint vector, and fusing the space-time joint vector, visual features, and semantic features to calculate a dynamic matching degree; Predicting the pet activity range based on the dynamic matching degree and generating a visualized heat map, and screening and sorting the dynamic matching degree based on the visualized heat map and a preset screening threshold, and outputting an interactive map interface.

[0005] In one embodiment, visual features of pet images and semantic features of text descriptions are extracted through a deep learning model, including: Using the Vision Transformer model, the pet image is segmented into multiple pixel blocks, and each pixel block is linearly projected to generate initial image features, and the initial image features are optimized based on a multi-level self-attention mechanism to generate optimized image features; Perform multi-instance weak supervision learning on the optimized image features, and use the MS-Captivator model to generate structured pet labels. The structured pet labels include breed classification codes, color spectrum vectors, body shape topologies, and special marker position matrices; Perform multi-head self-attention dynamic weighting on the structured pet labels, calculate the weights of each feature channel through the following formula, and output visual features: ; where Q is the query vector, is the key vector of the i-th feature channel, and d is the feature dimension; Input the text description into the RoBERTa model for semantic parsing to generate initial semantic features, and use the structured pet labels as cross-modal prior knowledge. Through the cross-modal cross-attention mechanism, the visual features are fused with the initial semantic features to generate enhanced semantic features; Based on the enhanced semantic features, calculate the difference loss between the color spectrum vector and the color description in the text description through a color semantic calibration network, and backpropagate to optimize the parameters of the MS-Captivator model and the RoBERTa model, and output semantic features strongly associated with the visual features.

[0006] In one embodiment, the lost spatio-temporal information is encoded as a spatio-temporal joint vector, including: Taking the lost position coordinates as the center point, calculate the Gaussian kernel encoding of the geographical coordinates through the following formula: ; where σ = 0.1 is the sensitivity coefficient, obtained through training with historical data, is the lost position coordinate; Calculate the attenuation factor according to the lost time through the following calculation formula: ; t is the lost time, and 48 is the reference attenuation period; Concatenate the Gaussian kernel encoding and the attenuation factor to generate a spatio-temporal feature vector, and perform normalization processing on the spatio-temporal feature vector to obtain a spatio-temporal joint vector.

[0007] In one embodiment, visual features, semantic features, and spatio-temporal joint vectors are fused to calculate a dynamic matching degree, including: Obtain candidate animal images, and extract features from the candidate animal images through a deep learning model to obtain the visual features of the candidate animal images; Calculate the cosine similarity between the visual features and the visual features of the candidate animal images; Calculate the semantic similarity between the semantic features and the text description of the candidate animal through the BERTScore algorithm; Weightedly fuse the spatio-temporal joint vector with the cosine similarity and the semantic similarity to obtain a weighted fusion result; Perform Sigmoid normalization on the weighted fusion result to obtain the dynamic matching degree, which is used to represent the matching probability between the candidate animal and the lost pet.

[0008] In one embodiment, predicting the pet's activity range based on the dynamic matching degree and generating a visualized heat map includes: Based on the visual features, semantic features, and spatio-temporal joint vectors, obtain the basic movement speed parameter and calculate the body size adjustment factor; Based on the behavior features in the semantic features, calculate the behavior activity intensity coefficient through natural language processing analysis; Combine the basic movement speed parameter, the body size adjustment factor, the behavior activity intensity coefficient, and the lost time to calculate the dynamic activity radius; Based on the dynamic activity radius, call the Gaode Map API to generate a dynamic gradient heat circle, with the lost position coordinates as the center and the dynamic activity radius as the radius; Mark the positions in the dynamic gradient heat circle with a dynamic matching degree higher than the preset threshold as red hot spots to generate a visualized heat map.

[0009] In one embodiment, the basic movement speed parameter is obtained through the following steps: Construct a database containing 100,000 pet movement trajectories, each pet movement trajectory is associated with visual feature labels, semantic behavior labels, and spatio-temporal environmental data, and the spatio-temporal environmental data includes the timestamp of the trajectory point and the Gaussian kernel encoding of the corresponding position; Train the spatio-temporal behavior dual-channel LSTM network model through the database to obtain a pet activity prediction model; Input the visual features, semantic features, and spatio-temporal joint vectors into the pet activity prediction model to obtain spatio-temporal behavior joint features; Compare the body size with the preset standard body size matrix of the same breed, calculate the body size difference ratio, and adjust the spatio-temporal behavior joint features based on the calculated body size difference ratio to generate a body size calibrated movement speed parameter; Statistically analyze the movement speed parameters after body shape calibration, take the median value of pets of the same breed, and obtain the basic movement speed parameters.

[0010] In one embodiment, the dynamic matching degree is screened and sorted based on the visual heat map and the preset screening threshold, and an interactive map interface is output, including: Compare the dynamic matching degree with the preset screening threshold, screen out the candidate positions with a dynamic matching degree higher than the preset screening threshold, and sort them from high to low according to the dynamic matching degree; Load the map base map through the Echarts engine or the Mapbox GL JS framework, and set the initial center point of the map base map based on the lost position coordinates; Overlay the visual heat map and the candidate positions on the map base map to generate an interactive map interface; Among them, based on the interactive map interface, clickable marker points are embedded in the candidate positions. The clickable marker points are used to display the pet feature comparison chart after clicking. The pet feature comparison chart includes a visual feature comparison matrix and a semantic feature comparison matrix; Integrate a time slider control in the interactive map interface. The time slider control is used to dynamically display the distribution of the dynamic matching degree at different time points.

[0011] In a second aspect, the present application also provides a pet searching geographic information matching system based on machine learning and natural language processing, including: A data acquisition module for acquiring the pet image, text description, and lost spatio-temporal information input by the user. The lost spatio-temporal information includes lost position coordinates and lost time; A feature extraction module for extracting the visual features of the pet image and the semantic features of the text description through a deep learning model. The visual features include breed, color, body shape, and special marks. The semantic features include breed entities, color modifiers, and behavior features; A pet searching matching module for encoding the lost spatio-temporal information into a spatio-temporal joint vector, and fusing the spatio-temporal joint vector, visual features, and semantic features to calculate the dynamic matching degree; A pet searching visualization module for predicting the pet activity range based on the dynamic matching degree and generating a visual heat map, and screening and sorting the dynamic matching degree based on the visual heat map and the preset screening threshold, and outputting an interactive map interface.

[0012] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes, it implements the steps in the first aspect.

[0013] Fourthly, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the first aspect are implemented.

[0014] The above pet-seeking geographic information matching method based on machine learning and natural language processing can comprehensively collect key information related to the lost pet by obtaining the pet image, text description, and lost time and space information input by the user, providing a comprehensive data basis for subsequent feature extraction and matching. Secondly, visual features such as the breed, color, body shape, and special markings of the pet image and semantic features such as breed entities, color modifiers, and behavior features of the text description can be extracted through a deep learning model. Compared with the traditional pet-seeking method that only relies on simple text descriptions, it can describe the features of the pet more comprehensively and meticulously, reducing the matching error caused by incomplete information and improving the accuracy of feature expression. Subsequently, the lost time and space information is encoded into a spatio-temporal joint vector and fused with the visual features and semantic features, adding spatio-temporal dimension constraints to the matching process and making the matching result more in line with the actual situation. Finally, the pet's activity range is predicted based on the dynamic matching degree and a visual heat map is generated, which can then intuitively display the areas where the pet may appear, providing clear direction guidance for pet-seeking personnel. And by screening and sorting the dynamic matching degree based on the visual heat map and a preset screening threshold, an interactive map interface is output, facilitating users to query and filter information according to their own needs and improving the efficiency of pet-seeking.

[0015] Compared with the traditional pet-seeking methods, this method is based on machine learning and natural language processing technologies, realizing the automated and precise matching and positioning of lost pet information. The comprehensive information acquisition and multi-dimensional feature extraction provide sufficient information for the matching process, enhancing the adaptability to different description methods and information sources in the pet-seeking process. The dynamic matching degree calculation that fuses spatio-temporal information and the visual result output further enhances the positioning accuracy of the lost pet's location and the accuracy of information matching, improving the success rate of pet-seeking and reducing the time and cost of pet-seeking. In addition, the interactive map interface provides a reliable operation platform for pet-seeking personnel, effectively improving the efficiency and quality of pet-seeking and providing strong support for retrieving lost pets. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1Flowchart of a pet - finding geographic information matching method based on machine learning and natural language processing provided for an exemplary embodiment of the present invention; Figure 2 Schematic structural diagram of a pet - finding geographic information matching system based on machine learning and natural language processing provided for an exemplary embodiment of the present invention. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0019] In one embodiment, as Figure 1 shown, a pet - finding geographic information matching method based on machine learning and natural language processing is provided. In this embodiment, an example is given where this method is applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps: S101: Obtain the pet image, text description, and lost time - space information input by the user. The lost time - space information includes the lost location coordinates and the lost time.

[0020] Specifically, the pet - related information input by the user can be obtained through a user terminal such as a mobile APP or a web page. Among them, the pet image can include key information such as the pet's face and body characteristics, and the text description of the lost pet is obtained, such as information about the pet's habits, personality traits, and the state when it was lost. In addition, the lost time - space information of the lost pet needs to be obtained. The lost location coordinates can be obtained by means of the GPS positioning function of the mobile phone, or the user manually marks the specific location where the pet was lost on the map to obtain it, and the lost time can be accurate to the specific year, month, day, etc., for subsequent pet - finding processing.

[0021] S102: Extract the visual features of the pet image and the semantic features of the text description through a deep - learning model. The visual features include breed, color, body shape, and special marks. The semantic features include breed entities, color modifiers, and behavioral features.

[0022] Specifically, for pet images, a deep learning framework such as ViT, Vision Transformer, an extended form of the Transformer model, can be used. After splitting the pet images into multiple small patches, they are input into the model for processing. This model can automatically learn and extract visual features such as the breed, color, body size, and special markings of the pet. For example, through the training of a large number of pet images, this model can accurately identify that the currently lost pet is a Chinese rural dog, with yellow and white fur, a medium body size, and a black birthmark on its right hind leg. For text descriptions, pre-trained models based on the Transformer architecture such as BERT or RoBERTa can be used, combined with operations such as text preprocessing, part-of-speech tagging, and entity recognition in natural language processing technology, to extract semantic features such as breed entities, color modifiers, and behavioral characteristics. For example, for the text description "My white kitten is very lively and likes to chase the ball", through this model, "kitten" can be extracted as the breed entity, "white" as the color modifier, and "likes to chase the ball" as the behavioral characteristic.

[0023] S103: Encode the lost spatio-temporal information into a spatio-temporal joint vector, and fuse the spatio-temporal joint vector, visual features, and semantic features to calculate the dynamic matching degree.

[0024] Specifically, the lost position coordinates can be converted into vector form through a specific encoding method, and the lost time can be encoded into a vector according to relevant algorithms of the time series. Then, the two are combined into a spatio-temporal joint vector. And based on this spatio-temporal joint vector, visual features, and semantic features, corresponding weights can be assigned according to different types of features for fusion. For example, for some pets with obvious special markings, the weight of the visual features of the special markings can be set relatively high. Subsequently, preset matching algorithms such as cosine similarity calculation and machine learning model prediction can be used to calculate the dynamic matching degree according to the fused features. This dynamic matching degree can reflect the similarity between the currently obtained information and other pet information in the database.

[0025] S104: Predict the pet's activity range based on the dynamic matching degree and generate a visualized heat map, and filter and sort the dynamic matching degree based on the visualized heat map and a preset filtering threshold, and output an interactive map interface.

[0026] Specifically, based on the calculated dynamic matching degree, the possible activity range of the pet can be predicted by combining the pet's habits, the environment of the lost location, and time factors, using relevant algorithm models. And a visual heat map can be generated by calling tools such as Amap API and Echarts. In this visual heat map, the darker the color of the area, the higher the possibility of the pet appearing. In addition, by setting a preset screening threshold, the information with a dynamic matching degree higher than the threshold can be screened, and sorted from high to low according to the dynamic matching degree, and an interactive map interface is output. Through this interactive map interface, users can not only intuitively see the matching results on the map, but also click on the marked points on the map to obtain detailed pet information, such as the time and location of finding the pet, the description of pet characteristics, etc., which is convenient for users to further confirm whether it is the pet they lost, greatly improving the efficiency and accuracy of pet searching.

[0027] In the above method, by obtaining pet images, text descriptions, and lost spatio-temporal information, the visual, semantic, and spatio-temporal features can be comprehensively integrated, solving the problem of incomplete information caused by a single data type in traditional pet searching methods. Secondly, by using a deep learning model to extract the visual features and semantic features of the lost pet, the multi-dimensional capture of pet features is realized, which not only improves the richness of features, but also avoids the limitations of single-modal data in describing pet features. And by encoding the lost spatio-temporal information into a spatio-temporal joint vector and fusing it with the visual features and semantic features, the time and space rules of pet loss can be captured, making the calculated dynamic matching degree closer to the actual situation. Finally, predicting the pet's activity range based on the dynamic matching degree and generating a visual heat map, screening and sorting the dynamic matching degree based on a preset screening threshold, and outputting an interactive map interface, can not only intuitively view the possible activity areas of the pet, but also quickly locate the high-probability pet searching areas, reducing the time cost for users to search for information in a large amount of data, and improving the accuracy and practicality of pet searching.

[0028] In an exemplary embodiment, the visual features of the pet image and the semantic features of the text description are extracted by a deep learning model. The visual features include breed, color, body type, and special marks. The semantic features include breed entity, color modifier, and behavior features, including: Using the Vision Transformer model, the pet image is segmented into multiple pixel blocks, and each pixel block is linearly projected to generate initial image features, and the initial image features are optimized based on a multi-level self-attention mechanism to generate optimized image features; Performing multi-instance weak supervision learning on the optimized image features, and using the MS-Captivator model to generate structured pet labels. The structured pet labels include breed classification codes, color spectrum vectors, body type topology maps, and special mark position matrices; Perform multi-head self-attention dynamic weighting on the structured pet label, calculate the weights of each feature channel through the following formula, and output visual features: ; where Q is the query vector, is the key vector of the i-th feature channel, and d is the feature dimension; Input the text description into the RoBERTa model for semantic parsing to generate initial semantic features, and use the structured pet label as cross-modal prior knowledge to fuse the visual features and the initial semantic features through a cross-modal cross-attention mechanism to generate enhanced semantic features; Based on the enhanced semantic features, calculate the difference loss between the color spectrum vector and the color description in the text description through a color semantic calibration network, and backpropagate to optimize the parameters of the MS-Captivator model and the RoBERTa model, and output semantic features strongly associated with the visual features.

[0029] Specifically, when the Vision Transformer model processes an image, different from the traditional convolutional neural network that extracts local features by sliding convolution based on local receptive fields, it divides the image into multiple non-overlapping pixel blocks of a fixed size for processing. Secondly, a linear projection operation can be performed on each segmented pixel block to convert it into a low-dimensional feature vector. This initial image feature contains the local feature information of each pixel block. Based on this initial image feature, through a multi-level self-attention mechanism, when processing each pixel block, the information of other pixel blocks in the image is considered, so as to capture the long-range dependencies in the image. For example, when processing an image of a pet cat, the model not only focuses on the pixel blocks of the face, but can also associate with the pixel block information of the tail, limbs and other parts through the self-attention mechanism. After processing each pixel block and comprehensively considering the global information, an optimized image feature is generated. Subsequently, the obtained optimized image feature is subjected to multi-instance weak supervision learning. Multi-instance weak supervision learning means that in the case where the training data annotation is not very accurate, multiple related instances can be combined into a package for learning. In this embodiment, the MS-Captivator model is used to process the optimized image feature to generate a structured pet label. Among them, the breed classification code digitizes the breed information of the pet for subsequent model processing and recognition; the color spectrum vector models in the Lab or HSV color space to quantify the distribution, hue and saturation information of the main colors in the pet image, which is used to enhance the discriminability of color semantics, so as to accurately describe the hue, saturation and other characteristics of the pet color; the body shape topology map describes the body shape contour and proportion and other information of the pet through a certain algorithm and data structure; the special marker position matrix records the specific position information of special markers such as birthmarks and spots on the pet in the image.

[0030] Specifically, the multi-head self-attention mechanism is an extension of the self-attention mechanism. It can perform parallel calculations through multiple different attention heads to capture the relationships between features from different perspectives, enabling more comprehensive extraction of feature information. Therefore, the above formula can be used to dynamically adjust the importance of different feature channels such as breed classification codes, color spectrum vectors, body type topology maps, and special marker position matrices in generating the final visual features through multi-head self-attention, and then output the weighted visual features to highlight more critical pet features. And the text description input by the user is input into the RoBERTa model. Among them, RoBERTa is a pre-trained language model based on the Transformer architecture. Through pre-training on a large amount of text data, it can learn rich language knowledge and semantic representation capabilities. The model performs semantic parsing on the text description, including operations such as word segmentation, part-of-speech tagging, and syntactic analysis, to generate initial semantic features. Since images and texts belong to different modalities of information, the structured pet label, as a structured representation of image features, can provide additional image-related knowledge for the generation of text semantic features. Therefore, the above-generated structured pet label can be used as cross-modal prior knowledge. Through the cross-modal cross-attention mechanism, the model can dynamically allocate weights for fusion according to the degree of association between visual features and initial semantic features, generating enhanced semantic features to make the semantic information more comprehensive and accurate. In addition, a color semantic calibration network can be introduced. By comparing the color spectrum vector in the structured pet label with the color description content in the text description, the difference between the two is calculated and converted into a loss value. Subsequently, through the backpropagation algorithm, this loss value is fed back to the MS-Captivator model and the RoBERTa model, enabling the subsequent models to more accurately associate visual features and semantic features when processing pet images and text descriptions. Finally, semantic features strongly associated with visual features are output, ensuring that the features extracted from images and texts can complement each other, thereby improving the accuracy and reliability of the entire feature extraction process and providing high-quality feature data for subsequent tasks such as lost pet geographical information matching.

[0031] In an exemplary embodiment, encoding the lost spatio-temporal information into a spatio-temporal joint vector includes: Taking the lost position coordinates as the center point, calculating the Gaussian kernel encoding of geographical coordinates through the following formula: ; where σ = 0.1 is the sensitivity coefficient, obtained through training with historical data, is the lost position coordinates; Calculating the decay factor according to the lost time through the following calculation formula: ; t is the loss time, and 48 is the reference decay period; Concatenate the Gaussian kernel encoding and the decay factor to generate a spatio-temporal feature vector, and normalize the spatio-temporal feature vector to obtain a spatio-temporal joint vector.

[0032] Specifically, the Gaussian kernel function can map geographical coordinates to a continuous feature space. Therefore, with the lost position coordinates as the center point, through Gaussian kernel encoding, the similarity between geographical coordinates can be captured, making similar geographical coordinates have a higher similarity in the feature space. The sensitivity coefficient σ in the above formula is used to control the width of the Gaussian kernel function and determines the similarity threshold between geographical coordinates. A smaller σ value indicates higher sensitivity to geographical coordinates, while a larger σ value indicates lower sensitivity. And the time decay factor can be calculated according to the loss time to reflect the impact of time on the pet's activity range. Among them, the time decay factor reflects the impact of the loss time on the pet's activity range. As time goes by, the pet's activity range will gradually expand, so the matching weight will gradually decrease. Through this time decay factor, the weight of the matching result can be dynamically adjusted to make the matching result more in line with the actual situation. The reference decay period is an empirical value representing the reference period of time decay. In this embodiment, the reference decay period is set to 48 hours, indicating that the impact of time decay is more significant within 48 hours. Subsequently, the Gaussian kernel encoding and the time decay factor can be concatenated to generate a spatio-temporal feature vector, and the spatio-temporal feature vector is normalized to ensure the rationality and accuracy of the spatio-temporal feature vector, thereby improving the accuracy and efficiency of pet searching, and finally obtaining a spatio-temporal joint vector.

[0033] In an exemplary embodiment, fuse the visual feature, semantic feature and spatio-temporal joint vector to calculate the dynamic matching degree, including: Obtain a candidate animal image, and extract features from the candidate animal image through a deep learning model to obtain the visual feature of the candidate animal image; Calculate the cosine similarity between the visual feature and the visual feature of the candidate animal image; Calculate the semantic similarity between the semantic feature and the text description of the candidate animal through the BERTScore algorithm; Fuse the spatio-temporal joint vector, cosine similarity and semantic similarity with weights to obtain a weighted fusion result; Perform Sigmoid normalization on the weighted fusion result to obtain the dynamic matching degree, which is used to represent the matching probability between the candidate animal and the lost pet.

[0034] Specifically, cosine similarity measures the similarity between two vectors by calculating the cosine value of the angle between them. The cosine similarity between the visual features of the lost pet and those of the candidate animal image can be calculated according to the cosine similarity formula. The closer the similarity value is to 1, the more similar the directions of the two visual feature vectors are, that is, the higher the similarity degree between the lost pet and the candidate animal in appearance. The BERTScore algorithm is built based on the pre-trained BERT model (Bidirectional Encoder Representations from Transformers, BERT). The BERT model can be pre-trained on large-scale text data to learn rich language semantic representations. That is, the semantic features of the lost pet and the text description of the candidate animal can be respectively input into the preset BERT model, and then the semantics of the words, the grammar structure of the sentences, and the context of the text can be analyzed. Subsequently, the BERTScore algorithm can calculate the semantic similarity by comparing the distances between the two texts in the semantic representation space output by the model. The semantic similarity value reflects the closeness of the semantics of the two, and the higher the value, the more similar the semantics. Since visual features, semantic features, and spatio-temporal information all play important roles in measuring the matching degree between the candidate animal and the lost pet, but their importance may be different, they can be fused by weighted fusion according to preset weights to obtain a weighted fusion result. And in order to convert it into a numerical value that can intuitively represent the matching probability between the candidate animal and the lost pet, the Sigmoid function can be used for normalization to output a dynamic matching degree. The dynamic matching degree represents the matching probability between the candidate animal and the lost pet.

[0035] In an exemplary embodiment, predicting the pet's activity range based on the dynamic matching degree and generating a visual heat map includes: Based on the visual features, semantic features, and spatio-temporal joint vector, obtain the basic movement speed parameter and calculate the body size adjustment factor; Based on the behavioral features in the semantic features, calculate the behavioral activity intensity coefficient through natural language processing analysis; Combine the basic movement speed parameter, body size adjustment factor, behavioral activity intensity coefficient, and the lost time to calculate the dynamic activity radius; Based on the dynamic activity radius, call the Gaode Map API to generate a dynamic gradient heat circle, where the dynamic gradient heat circle is centered on the lost position coordinates and has a radius of the dynamic activity radius; Mark the positions in the dynamic gradient heat circle with a dynamic matching degree higher than the preset threshold as red hot spots to generate a visual heat map.

[0036] Specifically, the basic movement speed parameters of different pet breeds can be preset. For example, the basic movement speed of small dogs is 2 kilometers per hour, medium-sized dogs is 3 kilometers per hour, and large dogs is 4 kilometers per hour. These basic movement speed parameters reflect the average movement speed of different breeds of pets and provide basic data for subsequent prediction of the activity range. Moreover, according to the standard body size information of pets of the same breed and the body size information in the visual features, a body size adjustment factor can be calculated to adjust the basic movement speed, making the prediction results more in line with the actual situation. In addition, a pre-trained model based on Transformer can be used to perform dependency syntactic analysis on the behavioral features in the text description, identify the subject-predicate relationship, verb-object relationship, etc. in the text, extract the semantic information of the behavioral features, and calculate the behavioral activity intensity coefficient. This behavioral activity intensity coefficient reflects the behavioral activity level of the pet. Pets with higher activity levels usually have a larger activity range, so the behavioral activity intensity coefficient can be increased accordingly. Finally, the basic movement speed parameter, body size adjustment factor, behavioral activity intensity coefficient, and loss time can be combined to calculate the dynamic activity radius. This dynamic activity radius reflects the possible activity range of the pet during the loss time, and as time goes by and the behavioral features of the pet change, the activity radius will gradually expand.

[0037] Furthermore, the Gaode Map API can be called to generate a dynamic gradient heat map with the lost location coordinates as the center and the dynamic activity radius as the radius. This dynamic gradient heat map can reflect the probability of the pet's appearance through the depth of color, with darker colors indicating a higher probability. And the positions with a dynamic matching degree higher than the preset threshold in the dynamic gradient heat map can be marked as red hot spots. These red hot spots can represent the areas where the probability of the pet's appearance is relatively high and the possibility of matching the lost pet is relatively large after considering the visual features, semantic features, spatio-temporal information, and movement-related factors of the lost pet. Finally, by integrating the marked red hot spots with the dynamic gradient heat map, a visual heat map can be obtained. By viewing this map, users can intuitively understand the high-probability areas where the lost pet may appear, further improving the efficiency and accuracy of pet searching.

[0038] In an exemplary embodiment, the basic movement speed parameter is obtained through the following steps: Construct a database containing 100,000 pet movement trajectories, with each pet movement trajectory associated with visual feature tags, semantic behavior tags, and spatio-temporal environment data. The spatio-temporal environment data includes the timestamp of the trajectory point and the Gaussian kernel encoding of the corresponding location; Train the spatio-temporal behavior dual-channel LSTM network model through the database to obtain a pet activity prediction model; Input the visual features, semantic features, and spatio-temporal joint vector into the pet activity prediction model to obtain spatio-temporal behavior joint features; Compare the body shape with the preset standard body shape matrix of the same breed, calculate the body shape difference ratio, and adjust the spatio-temporal-behavior joint features based on the calculated body shape difference ratio to generate the moving speed parameter after body shape calibration; Conduct statistical analysis on the moving speed parameter after body shape calibration, and take the median value of pets of the same breed to obtain the basic moving speed parameter.

[0039] Specifically, the visual feature labels in the pet's movement trajectory can record in detail the appearance information such as the breed, color, body shape, and special marks of the pet. The semantic behavior labels can describe the semantic information such as the pet's behavior habits and personality traits, which helps to understand the pet's behavior habits. The spatio-temporal environmental data can effectively represent the geographical location of the pet and the information of its surrounding environment. Subsequently, the constructed database can be used to train the spatio-temporal-behavior dual-channel LSTM (Long Short-Term Memory) network model. Among them, the spatio-temporal-behavior dual-channel LSTM network model has two channels. One channel focuses on processing spatio-temporal environmental data and analyzing the changing rules of the pet's movement trajectory over time and space; the other channel focuses on processing semantic behavior labels and visual feature labels to explore the correlations between the visual features, behavior habits, and movement patterns of different pets. By continuously inputting a large amount of pet movement trajectory data into the model for training, the model can learn the activity rules of pets under different visual features, semantic behaviors, and spatio-temporal environmental conditions, and finally obtain a pet activity prediction model that can accurately predict the activities of pets. Input the visual features, semantic features, and spatio-temporal joint vector of the currently lost pet into the trained pet activity prediction model. The model can deeply analyze and fuse these input information, and extract the spatio-temporal-behavior joint features that can comprehensively reflect the spatio-temporal activity characteristics and behavior patterns of the pet. This joint feature contains information such as the possible movement trends of the pet based on its appearance and behavior habits under specific spatio-temporal conditions.

[0040] Further, compare the size of the lost pet with a preset standard size matrix for the same breed. This preset standard size matrix for the same breed is obtained through statistical analysis of the size data of a large number of pets of the same breed and has a certain degree of representativeness and standardization. Through comparison, the difference ratio between the current pet's size and the standard size can be calculated. Since size differences can affect the movement speed of pets, for example, pets with a relatively large size and deviating from the standard size may have a relatively slow movement speed. Therefore, based on this size difference ratio, the spatio-temporal-behavior joint features can be adjusted to more accurately reflect the impact of the pet's size factor on the movement speed, and then generate the movement speed parameters after size calibration. And for pets of the same breed, the median of their movement speed parameters can be calculated. Among them, the median is the value located in the middle position after sorting a set of data from small to large. Compared with the average value, it is more resistant to the interference of outliers and has better robustness. Therefore, by taking the median value of the movement speed parameters of pets of the same breed, the basic movement speed parameters can be obtained. This parameter serves as the basic reference value for subsequent calculation of pet movement-related indicators and provides a reliable basis for accurately predicting the pet's activity range.

[0041] In an exemplary embodiment, based on the visual heat map and a preset screening threshold, screen and sort the dynamic matching degrees, and output an interactive map interface, including: Compare the dynamic matching degrees with the preset screening threshold, screen out the candidate positions with dynamic matching degrees higher than the preset screening threshold, and sort them from high to low according to the dynamic matching degrees; Load the map base map through the Echarts engine or the Mapbox GL JS framework, and set the initial center point of the map base map based on the lost position coordinates; Overlay the visual heat map and the candidate positions on the map base map to generate an interactive map interface; Among them, based on the interactive map interface, clickable marker points are embedded in the candidate positions. The clickable marker points are used to display the pet feature comparison map after clicking. The pet feature comparison map includes a visual feature comparison matrix and a semantic feature comparison matrix; Integrate a time slider control in the interactive map interface. The time slider control is used to dynamically display the distribution of dynamic matching degrees at different time points.

[0042] Specifically, the calculated dynamic matching degree can be compared with a preset screening threshold, such as 0.8, to screen out candidate positions with a dynamic matching degree higher than the preset screening threshold. Then, the screened candidate positions are sorted in descending order of the dynamic matching degree to ensure that the user can view the position with the highest matching degree first, improving the efficiency of pet searching. Subsequently, the map base map can be loaded through the Echarts engine or the Mapbox GL JS framework. Among them, the Echarts engine can provide rich data visualization functions, while the Mapbox GL JS framework can provide efficient rendering and interaction functions for geographical data. Then, the initial center point of the map base map can be set according to the lost position coordinates to ensure that the initial display area of the map is related to the pet lost position. The generated visual heat map is overlaid on the map base map to generate an interactive map interface. And in the interactive map interface, clickable marker points can be embedded in the candidate positions. After the user clicks on the marker point, a pet feature comparison chart can be displayed. The pet feature comparison chart includes a visual feature comparison matrix and a semantic feature comparison matrix. Among them, the visual feature comparison matrix can display the comparison of the breed, color, body shape, and special marks of the pet; the semantic feature comparison matrix can display the comparison of the breed entity, color modifier, and behavior characteristics of the pet, which helps the user to compare the characteristics of the lost pet and the found pet in detail. In addition, the interactive map interface can integrate a time slider control to dynamically display the distribution of the dynamic matching degree at different time points. Schematically, the time slider control allows the user to select a specific time point and update the heat map and candidate positions according to the selected time point, dynamically displaying the matching degree score to further assist in searching for the lost pet.

[0043] Based on the same inventive concept, as Figure 2 shown, the embodiment of the present application also provides a pet searching geographical information matching system 200 based on machine learning and natural language processing. The implementation solution provided by this system to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more of the following system embodiments of the pet searching geographical information matching system based on machine learning and natural language processing can refer to the limitations on the pet searching geographical information matching method based on machine learning and natural language processing in the above text, and will not be repeated here. The system includes: A data acquisition module 201, configured to acquire the pet image, text description, and lost spatio-temporal information input by the user, where the lost spatio-temporal information includes lost position coordinates and lost time; A feature extraction module 202, configured to extract the visual features of the pet image and the semantic features of the text description through a deep learning model, where the visual features include breed, color, body shape, and special marks, and the semantic features include breed entity, color modifier, and behavior characteristics; The pet searching matching module 203 is used to encode the lost spatio-temporal information into a spatio-temporal joint vector, fuse the spatio-temporal joint vector, visual features, and semantic features, and calculate the dynamic matching degree; The pet searching visualization module 204 is used to predict the pet's activity range according to the dynamic matching degree and generate a visualized heat map, and filter and sort the dynamic matching degree based on the visualized heat map and a preset filtering threshold, and output an interactive map interface.

[0044] In the above system, the data acquisition module 201 can acquire the pet image, text description, and lost spatio-temporal information input by the user, effectively solving the problems of incomplete and inaccurate information acquisition in traditional pet searching methods, and comprehensively integrating the key information related to the lost pet, thus providing a comprehensive data basis for subsequent pet searching analysis. The feature extraction module 202 can extract the visual features of the pet image and the semantic features of the text description through a deep learning model, and can describe the features of the pet from multiple dimensions, avoiding the limitation that traditional methods are difficult to accurately identify the pet only relying on simple descriptions. The pet searching matching module 203 can encode the lost spatio-temporal information into a spatio-temporal joint vector, and fuse the visual features and semantic features through the lost time and location information to calculate the dynamic matching degree, which not only effectively solves the problem that traditional pet searching matching ignores spatio-temporal factors, but also can more scientifically evaluate the matching possibility, making the matching result more in line with the actual situation and improving the accuracy and reliability of the matching. The pet searching visualization module 204 can predict the pet's activity range according to the dynamic matching degree to generate a visualized heat map, and filter and sort the dynamic matching degree through a preset filtering threshold, and output an interactive map interface. This module not only effectively solves the problems of unintuitive display and poor interactivity of traditional pet searching results, but also can intuitively present the areas where the pet may appear, helping users to query and filter according to their own needs and improving the efficiency of pet searching.

[0045] In an exemplary embodiment, the present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the pet searching geographic information matching method based on machine learning and natural language processing of the present application are implemented. A multi-core processor is preferably used to improve the parallel processing ability of the system. Memory: Provide sufficient temporary storage space to support the operation of the program and the processing of data. The memory capacity should be large enough to accommodate a large amount of supply information and computing tasks.

[0046] In an exemplary embodiment, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the pet-finding geographic information matching method based on machine learning and natural language processing of the present application are implemented. The computer-readable storage medium may include: read-only memory (ROM, ReadOnly Memory), random access memory (RAM, Random Access Memory), solid state drives (SSD, SolidState Drives), or optical discs, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance Random Access Memory) and dynamic random access memory (DRAM, Dynamic RandomAccess Memory).

[0047] The above-described embodiments merely represent several implementation manners of the embodiments of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. A pet - seeking geographical information matching method based on machine learning and natural language processing, characterized in that, The method includes: Obtaining a pet image, a text description, and lost spatio-temporal information input by a user, where the lost spatio-temporal information includes lost position coordinates and a lost time; Extracting visual features of the pet image and semantic features of the text description through a deep learning model, where the visual features include breed, color, body type, and special markings, and the semantic features include breed entities, color modifiers, and behavioral features; Encoding the lost spatio-temporal information into a spatio-temporal joint vector, and fusing the spatio-temporal joint vector, the visual features, and the semantic features to calculate a dynamic matching degree; Predicting a pet activity range according to the dynamic matching degree and generating a visualized heat map, and screening and sorting the dynamic matching degree based on the visualized heat map and a preset screening threshold, and outputting an interactive map interface.

2. The method according to claim 1, characterized in that, The extracting visual features of the pet image and semantic features of the text description through a deep learning model includes: Segmenting the pet image into multiple pixel blocks, inputting each pixel block into a Vision Transformer model for processing, and generating initial image features through a multi-level self-attention mechanism; Performing multi-instance weak supervision learning on the initial image features, and using an MS-Captivator model to generate a structured pet label, where the structured pet label includes a breed classification code, a color spectrum vector, a body type topology map, and a special marking position matrix; Performing multi-head self-attention dynamic weighting on the structured pet label, calculating the weights of each feature channel through the following formula, and outputting the visual features: ; where Q is the query vector, is the key vector of the i-th feature channel, and d is the feature dimension; Inputting the text description into a RoBERTa model for semantic parsing to generate initial semantic features, and using the structured pet label as cross-modal prior knowledge, and fusing the visual features and the initial semantic features through an attention mechanism to generate enhanced semantic features; Based on the enhanced semantic features, calculating the difference loss between the color spectrum vector and the color description in the text description through a color semantic calibration network, and backpropagating to optimize the parameters of the MS-Captivator model and the RoBERTa model, and outputting the semantic features strongly associated with the visual features.

3. The method according to claim 1, characterized in that, The encoding the lost spatio-temporal information into a spatio-temporal joint vector includes: Taking the lost position coordinates as the center point, and calculating the Gaussian kernel encoding of the geographical coordinates through the following formula: ; where σ = 0.1 is the sensitivity coefficient, obtained through training with historical data, is the coordinate of the lost position; Calculating a decay factor according to the lost time through the following calculation formula: ; t is the lost time, and 48 is the reference decay period; Concatenating the Gaussian kernel encoding and the decay factor to generate a spatio-temporal feature vector, and performing normalization processing on the spatio-temporal feature vector to obtain the spatio-temporal joint vector.

4. The method according to claim 1, characterized in that, The fusing the visual features, the semantic features, and the spatio-temporal joint vector to calculate the dynamic matching degree includes: Calculating the cosine similarity between the visual features and the visual features of a candidate animal image; Calculating the semantic similarity between the semantic features and the text description of a candidate animal through the BERTScore algorithm; Perform weighted fusion on the spatio-temporal joint vector, the cosine similarity, and the semantic similarity to obtain a weighted fusion result; Perform Sigmoid normalization on the weighted fusion result to obtain the dynamic matching degree, which is used to represent the matching probability between the candidate animal and the lost pet.

5. The method according to claim 1, characterized in that, The predicting the pet activity range and generating a visual heat map according to the dynamic matching degree includes: According to the visual feature, the semantic feature, and the spatio-temporal joint vector, load the basic movement speed parameter and calculate the body size adjustment factor; Based on the behavior feature in the semantic feature, calculate the behavior activity intensity coefficient through dependency syntactic analysis; Combine the basic movement speed parameter, the body size adjustment factor, the behavior activity intensity coefficient, and the lost time to calculate the dynamic activity radius; Based on the dynamic activity radius, call the Gaode Map API to generate a dynamic gradient heat circle, where the dynamic gradient heat circle takes the lost position coordinates as the center and the dynamic activity radius as the radius; Mark the positions in the dynamic gradient heat circle where the dynamic matching degree is higher than the preset threshold as red hot spots to generate the visual heat map.

6. The method according to claim 5, characterized in that, The obtaining the basic movement speed parameter through the following steps: Construct a database containing 100,000 pet movement trajectories, each of which is associated with a visual feature label, a semantic behavior label, and spatio-temporal environment data, and the spatio-temporal environment data includes the timestamp of the trajectory point and the Gaussian kernel encoding of the corresponding position; Train the spatio-temporal behavior dual-channel LSTM network model through the database to obtain a pet activity prediction model; Input the visual feature, the semantic feature, and the spatio-temporal joint vector into the pet activity prediction model to obtain a spatio-temporal behavior joint feature; Compare the body size with the preset standard body size matrix of the same breed, calculate the body size difference ratio, and adjust the spatio-temporal behavior joint feature based on the calculated body size difference ratio to generate a movement speed parameter after body size calibration; Perform statistical analysis on the movement speed parameter after body size calibration, and take the median value of pets of the same breed to obtain the basic movement speed parameter.

7. The method according to claim 1, wherein, The screening and sorting the dynamic matching degree based on the visual heat map and a preset screening threshold, and outputting an interactive map interface includes: Compare the dynamic matching degree with the preset screening threshold, screen out the candidate positions where the dynamic matching degree is higher than the preset screening threshold, and sort them from high to low according to the dynamic matching degree; Load the map base map through the Echarts engine or the Mapbox GL JS framework, and set the initial center point of the map base map based on the lost position coordinates; Overlay the visual heat map and the candidate positions on the map base map to generate the interactive map interface; Among them, based on the interactive map interface, clickable marker points are embedded in the candidate positions, and the clickable marker points are used to display a pet feature comparison map after clicking, and the pet feature comparison map includes a visual feature comparison matrix and a semantic feature comparison matrix; Integrate a time slider control in the interactive map interface, where the time slider control is used to dynamically display the dynamic matching degree distribution at different time points.

8. A pet-searching geographic information matching system based on machine learning and natural language processing, wherein, The system includes: A data acquisition module, configured to acquire a pet image, a text description, and lost space-time information input by a user, where the lost space-time information includes lost position coordinates and a lost time; A feature extraction module, configured to extract visual features of the pet image and semantic features of the text description through a deep learning model, where the visual features include breed, color, body type, and special markings, and the semantic features include breed entities, color modifiers, and behavioral features; A pet searching and matching module, configured to encode the lost space-time information into a space-time joint vector, and fuse the space-time joint vector, the visual features, and the semantic features to calculate a dynamic matching degree; A pet searching and visualization module, configured to predict a pet activity range according to the dynamic matching degree and generate a visualized heat map, and filter and sort the dynamic matching degree based on the visualized heat map and a preset filtering threshold, and output an interactive map interface.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, wherein, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Recommendation method and system based on pet feature tag, program product and medium

    CN113722582A

  • Vision-text collaborative abstract generation method and system based on multi-modal learning

    CN119862861A

  • My pet ID collar

    US20080133256A1

  • Ranked Insight Machine Learning Operation

    US20180232659A1

Cited By

  • Lost and found method and system based on intelligent bus platform

    CN120707888A

  • Image recognition method based on visual language, controller, robot and medium

    CN120726435A

  • Method, system, device and medium for dynamic visualization generation based on user behavior and context awareness

    CN121051169B

  • Panda field tracking method based on deep learning

    CN121074155A