Information processing method and related device

By matching probabilistic text embedding vectors with image features, the accuracy and efficiency issues of image scene recognition are solved, enabling automatic acquisition and labeling of camera location information and improving the automation level of video surveillance systems.

WO2026020684A9PCT designated stage Publication Date: 2026-03-26CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

The accuracy and efficiency of image scene recognition in existing technologies need to be improved. Video surveillance systems cannot automatically acquire and label camera location information, resulting in huge manpower input and large errors.

Method used

A method of matching probabilistic text embedding vectors with image features is adopted. By freezing the pre-trained text encoder and image encoder, probabilistic text embedding vectors are generated, and the matching degree is adjusted using Gaussian distribution to realize the recognition of target scene and automatic acquisition of camera position information.

Benefits of technology

It improves the accuracy and efficiency of image scene recognition, realizes the automatic acquisition and labeling of camera location information, and reduces manpower input and errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137855_26032026_PF_FP_ABST
    Figure CN2024137855_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an information processing method and a related device, which relate to the technical field of computers and communications. The method comprises: obtaining probabilistic text embedding vectors (S110); obtaining an image feature of an image, and obtaining the degrees of matching between the probabilistic text embedding vectors and the image feature (S120); using the probabilistic text embedding vector with the highest degree of matching as a target probabilistic text embedding vector (S130); and determining a target prompt template corresponding to the target probabilistic text embedding vector, and on the basis of target scenario prompt information corresponding to the target prompt template, determining a target scenario corresponding to the image (S140).
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method and related device

[0001] Cross-reference to Related Applications

[0002] The present disclosure claims priority to Chinese Patent Application No. 202410993678.8, filed on July 23, 2024, entitled “Information processing method and related device”, the entire contents of which are incorporated herein by reference in its entirety. TECHNICAL FIELD

[0003] The present disclosure relates to the field of computer and communication technology, and in particular, to an information processing method, an information processing apparatus, an electronic device, a computer readable storage medium, and a computer program product. BACKGROUND

[0004] In some application scenarios, it is necessary to determine the scene corresponding to an image by processing and recognizing the image. The solutions provided in the related art need to further improve the accuracy and efficiency of image scene recognition. SUMMARY

[0005] An information processing method is provided in an embodiment of the present disclosure. The method includes obtaining a probability text embedding vector, obtaining an image feature of an image, and obtaining a matching degree between the probability text embedding vector and the image feature, taking a probability text embedding vector with the largest matching degree as a target probability text embedding vector, determining a target prompt template corresponding to the target probability text embedding vector, and determining a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template. In an exemplary embodiment, the method can be executed by an electronic device (such as a terminal and / or a server), or the method is executed by an apparatus (such as a chip, etc.) configured in the electronic device.

[0006] An information processing apparatus is provided in an embodiment of the present disclosure. In one design, the information processing apparatus can include a module corresponding to each of the methods / operations / steps / actions described in any embodiment of the present disclosure. The module can be a hardware circuit, a software, or a combination of hardware circuit and software. In one design, the information processing apparatus includes a processing unit configured to obtain a probability text embedding vector, obtain an image feature of an image, and obtain a matching degree between the probability text embedding vector and the image feature, take a probability text embedding vector with the largest matching degree as a target probability text embedding vector, determine a target prompt template corresponding to the target probability text embedding vector, and determine a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template.

[0007] The embodiment of the present disclosure further provides a processor, comprising: an input circuit, an output circuit and a processing circuit. The processing circuit is configured to receive a signal through the input circuit and transmit a signal through the output circuit, so that the processor executes the method in any embodiment of the present disclosure.

[0008] In the exemplary embodiment, the processor can be one or more chips, the input circuit can be an input pin, the output circuit can be an output pin, and the processing circuit can be a transistor, a gate circuit, a flip-flop, various logic circuits and the like. The input signal received by the input circuit can be received and input by, for example but not limited to, a receiver, the signal output by the output circuit can be output to and transmitted by, for example but not limited to, a transmitter, and the input circuit and the output circuit can be the same circuit which is used as the input circuit and the output circuit at different times. The embodiment of the present disclosure does not limit the specific implementation of the processor and various circuits.

[0009] The embodiment of the present disclosure provides an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the information processing method in any embodiment of the present disclosure via execution of the executable instructions.

[0010] The embodiment of the present disclosure provides a computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the information processing method in any embodiment of the present disclosure.

[0011] The embodiment of the present disclosure provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the information processing method in any embodiment of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0012] The drawings herein are incorporated into the specification and form a part of the specification, show embodiments consistent with the present disclosure, and together with the specification serve to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0013] FIG. 1 shows a flowchart of an information processing method in an embodiment of the present disclosure.

[0014] FIG. 2 shows a schematic diagram of an information processing method in an embodiment of the present disclosure.

[0015] FIG. 3 shows a schematic diagram of a probability prompt learning module in an embodiment of the present disclosure.

[0016] FIG. 4 shows a schematic diagram of another information processing method in an embodiment of the present disclosure.

[0017] FIG. 5 shows a structural block diagram of an information processing apparatus in an embodiment of the present disclosure.

[0018] FIG. 6 shows a structural block diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations.

[0020] In addition, the accompanying drawings are included to provide a thorough understanding of the present disclosure and are not intended to be exhaustive or to limit the present disclosure to the precise outlines described. The same or similar components referred to with the same or similar reference numerals in different drawings represent the same or similar parts, and thus repetitive description thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, and do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0021] Some of the terms involved in the embodiments of the present disclosure are explained below.

[0022] General large model: A large model can also be referred to as a Foundation Model or a Pre-Training Unit (PTU). The model extracts and learns knowledge from a large amount of corpus or image, and then produces a large model with a large number of parameters.

[0023] Prompt learning: In the case of not significantly changing the structure of the pre-training model, the parameters of the model are fine-tuned by adding "prompt information" to the input, changing the downstream task to a text generation task, and the like. For example, in the following example, K (K is a positive integer greater than or equal to 1, and hereinafter can also be referred to as the second number) prompt templates are provided, and it is determined which scene the image or sample image belongs to among the K prompt templates. When fine-tuning the parameters of the model, a loss function can be calculated based on the predicted scene and the real scene label, and the model parameters can be adjusted based on the loss function.

[0024] Probabilistic learning: how individuals learn to predict the occurrence or nonoccurrence of events under the principle of probability. It should be understood that although the following embodiments are exemplified by Gaussian distribution, the present disclosure is not limited thereto, and if other distributions are used, the following formulas can be changed accordingly, for example, a uniform distribution. Different distributions have different effects for different data sources, and some Gaussian distributions work better, while some uniform distributions work better. The actual scene can be selected accordingly.

[0025] Vision Internet: Vision Internet is the fifth basic network in addition to mobile network, broadband network, Internet of Things, and satellite network, which is a video backhaul and processing network.

[0026] The specific implementation of the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0027] FIG. 1 shows a flowchart of an information processing method according to an embodiment of the present disclosure. The method provided in the embodiment of FIG. 1 can be executed by any electronic device, such as a terminal and / or a server, and the present disclosure is not limited thereto. As shown in FIG. 1, the method provided in the embodiment of the present disclosure can include the following steps.

[0028] In S110, a probability text embedding vector is obtained.

[0029] In the embodiment of the present disclosure, the probability text embedding vector refers to a vector for representing the sampling probability of each prompt template. The greater the sampling probability of a prompt template indicated in the probability text embedding vector, the more likely the prompt template is selected for matching the image features / sample image features. In some embodiments, the dimension of the probability text embedding vector can be the same as the number of prompt templates. For example, if there are K (also referred to as the second number) prompt templates, each probability text embedding vector has K dimensions, where the value of the kth dimension indicates the probability of selecting the kth prompt template. Wherein, k is a positive integer greater than or equal to 1 and less than or equal to K.

[0030] In an exemplary embodiment, the method provided in the embodiment of the present disclosure further includes obtaining a second number of prompt templates; inputting the second number of prompt templates into the frozen and pre-trained text encoder to obtain a second number of text vectors.

[0031] In the embodiments of the present disclosure, the prompt template refers to embedding prompt information in a preset template, and the prompt information is used to assist or guide the model to make a prediction, so as to improve the accuracy and speed of the prediction. The preset template can be set according to the actual scene, and the specific form of the template is not limited in the present disclosure. The prompt information can be determined according to the actual required prediction output content, for example, if the scene of the image needs to be predicted, the prompt information can include scene prompt information. The scene prompt information refers to prompt information used to assist or guide the model to predict the scene corresponding to the image.

[0032] In the embodiments of the present disclosure, the text encoder refers to a machine learning model or a deep learning model that can perform encoding processing on text data. The type and structure of the text encoder are not limited in the embodiments of the present disclosure.

[0033] In the embodiments of the present disclosure, the text encoder can be pre-trained using a text data set. Before obtaining the probability text embedding vector, the K prompt templates can be input into the pre-trained text encoder, and in the process of obtaining the K text vectors corresponding to the K prompt templates, the text encoder is frozen, that is, the model parameters of the pre-trained text encoder are used at this time, and the model parameters of the text encoder are not updated or adjusted. The K text vectors can be used to generate the probability text embedding vector.

[0034] In the example embodiments, a first number of probability text embedding vectors are included, the probability distribution embedding vector is of a second number of dimensions, and each dimension indicates a sampling probability of a corresponding prompt template. Each prompt template includes corresponding scene prompt information. That is, C probability text embedding vectors are included, the probability distribution embedding vector is of K dimensions, and the kth dimension indicates a sampling probability of the kth prompt template in the K prompt templates. Wherein, K and C are positive integers greater than or equal to 1, K is greater than C, k is a positive integer greater than or equal to 1 and less than or equal to K, and the kth prompt template includes the kth scene prompt information.

[0035] In the embodiments of the present disclosure, C probability text embedding vectors are generated based on the K prompt templates, and each probability text embedding vector includes K dimensions, and the value of the kth dimension indicates the sampling probability of the kth prompt template. On the one hand, since the number of prompt templates in the actual scene is usually large, for example, K is equal to 50, 100, 200, etc., if K text vectors of the K prompt templates are directly matched with the image features, a large amount of computing resources will be occupied, and the matching speed and efficiency will be reduced, while the present disclosure uses smaller C probability text embedding vectors to match with the image features, which can reduce the consumption of computing resources and improve the matching speed and efficiency, for example, C can be equal to 5, 10 or 15, etc. On the other hand, by keeping the dimension of the probability distribution embedding vector consistent with the number of prompt templates, and letting the kth dimension of the probability distribution embedding vector indicate the sampling probability of the kth prompt template, the complete K prompt templates can be represented by the C probability text embedding vectors.

[0036] In the embodiments of the present disclosure, the kth prompt template includes kth scene prompt information, and the scene prompt information included in different prompt templates is different.

[0037] In the example embodiments, obtaining a probability text embedding vector includes: obtaining a first number of candidate probability text embedding vectors, the candidate probability distribution embedding vector has a second number of dimensions, each dimension indicates a candidate sampling probability of a corresponding prompt template (for example, the kth dimension indicates the candidate sampling probability of the kth prompt template); obtaining a sample image feature of a sample image and scene label information of the sample image; obtaining a sample matching degree between each candidate probability text embedding vector and the sample image feature respectively; taking the candidate probability text embedding vector with the largest sample matching degree as a target candidate probability text embedding vector; determining the prompt template corresponding to the target candidate probability text embedding vector and the scene prompt information corresponding thereto, to determine the scene prediction information corresponding to the sample image; obtaining a difference degree between the scene prediction information of the sample image and the scene label information; and adjusting the first number of candidate probability text embedding vectors according to the difference degree to obtain the first number of probability text embedding vectors.

[0038] In the embodiments of the present disclosure, the candidate probability text embedding vector refers to the probability text embedding vector in the adjustment process using the sample image and the corresponding scene label information. When the difference degree between the predicted scene prediction information of the sample image and the scene label information thereof satisfies a preset condition by using the candidate probability text embedding vector, the candidate probability text embedding vector at this time can be taken as the probability text embedding vector, that is, the probability text embedding vector is the candidate probability text embedding vector after adjustment using the sample image and the corresponding scene label information, and the difference degree satisfies the preset condition.

[0039] The scene label information of the sample image in the embodiments of the present disclosure refers to the pre-labeled, real scene of the sample image. The scene prediction information of the sample image refers to the scene of the sample image predicted based on the model output, which may be consistent with the real scene or may have a certain deviation, or even be wrong.

[0040] In the embodiments of the present disclosure, the difference degree refers to the deviation or index between the scene prediction information of the sample image and its real scene label information. There can be various ways to measure the difference degree, such as the distance (such as Euclidean distance, but the present disclosure is not limited thereto) between the scene prediction information and the scene label information, the loss function (such as cross-entropy loss function), etc., and the present disclosure does not limit this.

[0041] In the embodiments of the present disclosure, the preset condition can be set according to actual needs. For example, when the adjusted candidate probability text embedding vector is used for feature matching with the sample image feature, and the predicted scene prediction information of the sample image is completely consistent with the real scene label information (the difference degree is, for example, 0), it is considered that the difference degree meets the preset condition. However, the present disclosure is not limited thereto.

[0042] When the difference degree does not meet the preset condition, the values of each dimension in the candidate probability text embedding vector can be adjusted again, for example, another C prompt templates can be selected from the remaining K-C prompt templates. Here, the remaining K-C prompt templates refer to prompt templates other than the first randomly selected C prompt templates. If the difference degree still does not meet the preset condition after adjustment, the values of each dimension in the candidate probability text embedding vector are adjusted again, for example, another C prompt templates can be selected from the remaining K-2C prompt templates. Repeat the execution in turn until the difference degree meets the preset condition.

[0043] In an example embodiment, obtaining the first quantity of candidate probability text embedding vectors comprises: obtaining a first quantity of initial probability text embedding vectors, each initial probability text embedding vector being of a second quantity of dimensions, each dimension indicating an initial probability of a corresponding prompt template (e.g., the kth dimension indicating an initial probability of the kth prompt template); taking each initial probability text embedding vector as a center vector of a corresponding candidate probability text embedding vector; obtaining a diagonal covariance matrix of each initial probability text embedding vector; and determining a probability density function of each candidate probability text embedding vector based on the initial probability text embedding vector as a Gaussian probability distribution having the center vector and the diagonal covariance matrix, thereby obtaining the corresponding candidate probability text embedding vector. For example, taking the cth initial probability text embedding vector as a center vector of a cth candidate probability text embedding vector; obtaining a diagonal covariance matrix of the cth initial probability text embedding vector; and determining a probability density function of the cth candidate probability text embedding vector based on the cth initial probability text embedding vector as a Gaussian probability distribution having the center vector and the diagonal covariance matrix, thereby obtaining the cth candidate probability text embedding vector.

[0044] In an example embodiment, the candidate probability text embedding vectors are initialized to obtain the initial probability text embedding vectors. For example, during initialization, C prompt templates are randomly selected from the K prompt templates, and the C initial probability text embedding vectors are used to represent the selected C prompt templates. For example, for a certain initial probability text embedding vector, if it is represented as {1, 0, 0, …, 0}, it indicates that the 1st prompt template in the K prompt templates is selected. For another example, if a certain initial probability text embedding vector is represented as {0, 0, 1, …, 0}, it indicates that the 3rd prompt template in the K prompt templates is selected. That is, when a certain initial probability text embedding vector indicates that the kth prompt template is selected, the kth dimension of the initial probability text embedding vector has a value greater than the values of the other dimensions.

[0045] In an example embodiment, each initial probability text embedding vector in the C initial probability text embedding vectors is processed to satisfy a Gaussian distribution, so that the C candidate probability text embedding vectors obtained can represent candidate sampling probabilities of the K prompt templates. At the same time, it becomes regular to select C prompt templates from the K prompt templates for matching with image features or sample image features.

[0046] In an example embodiment, obtaining the first quantity of candidate probability text embedding vectors further comprises: resampling each candidate probability text embedding vector to satisfy a normal distribution.

[0047] Optionally, after each candidate probability text embedding vector is made to satisfy a Gaussian distribution, each candidate probability text embedding vector can also be resampled to satisfy a normal distribution. That is, the candidate probability text embedding vector satisfying the normal distribution can be used to perform feature matching with the sample image feature to obtain a sample matching degree.

[0048] In S120, an image feature of an image is obtained, and a matching degree between the probability text embedding vector and the image feature is obtained.

[0049] In the embodiments of the present disclosure, the matching degree refers to the similarity between the probability text embedding vector and the image feature. The higher the matching degree, the higher the similarity; the lower the matching degree, the lower the similarity.

[0050] In the example embodiments, obtaining a sample image feature of a sample image includes: inputting the sample image into a frozen and pre-trained image encoder to obtain the sample image feature.

[0051] The image encoder (Image encoder) in the embodiments of the present disclosure refers to a machine learning or deep learning model used to extract image features to realize encoding of information or data in an image. The present disclosure does not limit the type or structure of the image encoder. In the embodiments of the present disclosure, the image encoder is first pre-trained using an image data set. When obtaining a sample image feature, the sample image is input into the pre-trained image encoder, and the model parameters of the image encoder are frozen during this period, that is, the model parameters of the pre-trained image encoder are used.

[0052] In the example embodiments, the image feature includes K dimensions, and K is a positive integer greater than or equal to 1. The cross entropy between the probability text embedding vector and the image feature is obtained based on the following formula, including:

[0053] wherein c is a positive integer greater than or equal to 1 and less than or equal to C, C is the number of the number of probability text embeddings, and C is a positive integer greater than or equal to 1; q (d,c) represents the cross entropy between the cth probability text embedding vector and the image feature; represents the kth dimension of the image feature, and k is a positive integer greater than or equal to 1 and less than or equal to K; represents the kth dimension in the cth probability text embedding vector.

[0054] In some embodiments, the matching degree between the probability text embedding vector and the image feature can be measured by calculating the cross-entropy between each probability text embedding vector and the image feature. The matching degree calculated in this way is inversely related to the cross-entropy, that is, the greater the cross-entropy, the lower the matching degree; the smaller the cross-entropy, the higher the matching degree.

[0055] It can be understood that although the above-mentioned way of measuring the matching degree is proposed in the embodiments of the present disclosure, the present disclosure is not limited thereto, for example, the matching degree can also be measured by the cosine similarity, Euclidean distance, etc. between the probability text embedding vector and the image feature.

[0056] In some embodiments, the sample matching degree can be calculated with reference to formula (1). At this time, q (d,c) may represent the cross-entropy between the cth candidate probability text embedding vector and the sample image feature; may represent the kth dimension of the sample image feature. may represent the kth dimension of the cth candidate probability text embedding vector. However, the present disclosure is not limited thereto.

[0057] In S130, the probability text embedding vector with the largest matching degree is taken as the target probability text embedding vector.

[0058] For example, the probability text embedding vector with the largest matching degree can be selected from the C probability text embedding vectors as the target probability text embedding vector.

[0059] In S140, the target prompt template corresponding to the target probability text embedding vector is determined, and the target scene corresponding to the image is determined according to the target scene prompt information corresponding to the target prompt template.

[0060] In the embodiments of the present disclosure, according to the size of the value of each dimension in the K-dimensional target probability text embedding vector, the prompt template corresponding to the dimension with the largest value can be determined as the target prompt template. For example, assuming that the target probability text embedding vector is {0.7, 0.001, 0.002, 0, …, 0}, that is, the value of the 1st dimension in the K-dimensional vector is the largest, equal to 0.7, since the 1st dimension corresponds to the 1st prompt template in the K prompt templates, the 1st prompt template is determined as the target prompt template.

[0061] In the embodiments of the present disclosure, each prompt template in the K prompt templates includes corresponding scene prompt information. The scene prompt information in the target prompt template is taken as the target scene prompt information. For example, if the 1st prompt template is the target prompt template, and the target scene prompt information contained in the 1st prompt template is “kitchen”, the target scene prompt information is determined as “kitchen”, and thus the target scene corresponding to the image is determined as “kitchen”.

[0062] In an example embodiment, the method provided by the embodiments of the present disclosure further includes: acquiring the image from an image acquisition device; and determining point position information of the image acquisition device according to a target scene corresponding to the image.

[0063] In the embodiments of the present disclosure, an image can be acquired by an image acquisition device (for example, a camera or a monitoring camera, but the present disclosure is not limited thereto, and examples will be given below), and the image is sent to the electronic device. The electronic device can acquire the image and identification information of the image acquisition device, and the identification information is used to uniquely identify the image acquisition device, that is, to indicate that the image is collected by the image acquisition device. By processing the image by the above method, the target scene corresponding to the image can be determined, so that the target scene can be used as the point position information of the image acquisition device, that is, the point position information herein refers to the scene collected or located or monitored by the image acquisition device. Thus, the automatic acquisition and automatic labeling of the point position information of the image acquisition device can be realized.

[0064] The information processing method provided by the embodiments of the present disclosure performs feature matching based on a probability text embedding vector and an image feature to identify a target scene corresponding to an image, which can improve the efficiency and accuracy of image scene recognition.

[0065] The 5G (5th Generation Mobile Communication Technology) era has given rise to a large number of video-oriented applications, such as smart and safe cities, the Internet, unmanned driving, video monitoring, and other image and video-based content, which have very wide application scenarios. Image and video-related applications will become one of the main sources of incremental traffic in the 5G and post-5G eras.

[0066] Under the continuous empowerment and driving of new infrastructure and new technologies such as big data, AI (Artificial Intelligence), 5G, cloud computing, and IoT (Internet of Things), in line with the trend of digital services and industrial development, the video network (also known as video network) has been released.

[0067] With the rapid development of 5G, big data, and artificial intelligence, video has become ubiquitous in work and life. Video application scenarios, such as video monitoring scenarios, lack the point position information (or point position or point position information) of monitoring cameras (also referred to as cameras, which are a type of image acquisition device). For example, a building has 50 cameras, and the point position information of these 50 cameras, such as which cameras are located at the stairs, corridors, conference rooms, and kitchens, cannot be obtained.

[0068] The video surveillance system in the relevant technology cannot automatically acquire and label the location information of the cameras. It can only rely on human eyes to view the video and manually label the location information of the cameras. When the camera has historical tags, the manually labeled location information is manually compared with the historical tags of the camera to finally determine the location information of the camera.

[0069] The mainstream framework for acquiring camera location information in video networks in related technologies mainly includes the following modules:

[0070] (1) Image information input section, which is used to input the image information of the camera in the video network to the system and transmit it to the manual processing module section.

[0071] (2) Manual processing module, which manually processes the image information of the camera in the video network, and understands the scene information or scene label information in the camera image based on human eyes.

[0072] (3) Tag information comparison module compares the existing tag location information (i.e., historical tags or historical location information) of the camera with the scene tag information in the image as understood by the human eye. The historical location information was manually labeled before this.

[0073] (4) Obtain the label comparison results.

[0074] (5) Obtain the camera's location information / location information / location information based on the label comparison results.

[0075] This is because, in real-world scenarios, the location or scene of some cameras may change. For example, the place monitored by the camera might be renovated; previously a kitchen, it might become a barbershop. This leads to inconsistencies between the previously classified scene and the current classification. Due to the large number of cameras, manually updating the location information typically takes a long time, resulting in incorrect camera location information stored in the system, causing difficulties or errors in subsequent processing. However, the solution or large model provided in this disclosure can automatically and periodically scan the images captured by the cameras, perform image recognition processing, and update the camera location information in a timely manner. This improves the efficiency of updating camera location information and prevents errors in subsequent analysis and processing caused by incorrect location information.

[0076] The system modules in the related technologies have the following technical problems:

[0077] (1) The video network in related technologies is to understand the scene by viewing the camera images with the human eye. However, the video network has tens of millions of cameras, and it is operated manually by operators, which requires a huge investment of manpower and is labor-intensive and time-consuming.

[0078] (2) The video network in related technologies does not understand the content of the image / video itself, but compares the text information of the video tags. Some of the video tags are manually annotated in history, and due to the long time, there are certain errors.

[0079] (3) In the video surveillance system of related technologies, the tag information comparison module compares the existing tag location information with the scene tag information in the image as understood by the human eye. This is done manually by the operators, which is a huge workload.

[0080] As can be seen from the above, the video systems or video surveillance systems in the related technologies cannot automatically acquire the location information of each camera, resulting in the inability to automatically label the location information of each camera; nor can they understand the scene information of the cameras based on the keywords of the location information. In these related technologies, the only way is for humans to view the surveillance video and manually label the location information of each camera. The video network management platform in these technologies has approximately 30 million front-end cameras, making the viewing of video surveillance extremely time-consuming and labor-intensive. Furthermore, the existing artificial intelligence and deep learning algorithms in these technologies cannot meet the application requirements of accuracy, scene understanding, and abstract semantics for intelligent video analysis tasks.

[0081] In some scenarios, existing technologies cannot accurately locate camera positions based on keywords in location information. However, using the method provided in this disclosure, after obtaining the latitude and longitude (or geographical location information) and location information of the camera, the scene category carried by the latitude and longitude and location information can be combined to obtain more accurate camera positioning. For example, it can be distinguished that there are 4 cameras in the kitchen, 1 in the kitchen on the 6th floor, and 1 in the kitchen in the garden, thus locating the kitchens where the cameras are deployed in different locations.

[0082] The purpose of the method provided in this disclosure is to overcome the shortcomings of related technologies, and to achieve automatic acquisition of the location information of each camera; automatic annotation of the location information of each camera; and understanding of camera scene information based on keywords in the location information. Optionally, it can also perform high-accuracy camera location information positioning based on keywords in the location information.

[0083] The method provided by the embodiments of the present disclosure is designed for a large model parameter fine-tuning based on a probability distribution prompt learning method in a videoconferencing and especially a monitoring video system, and a cross entropy based pixel text matching loss calculation method (which can refer to the above formula (1)) is designed. When a camera deployed in a new place, for example, a new city, needs to be scene-identified, the accuracy of scene category judgment can be improved. The point information of each camera in the videoconferencing can also be automatically acquired and automatically labeled.

[0084] As shown in FIG. 2, the prompt template 11 is input to the text encoder 12 to obtain a text vector of the prompt template. The text encoder 12 is connected with the probability prompt learning module 13, so that the probability prompt learning module 13 can generate a candidate probability distribution embedding vector and a probability text embedding vector based on the prompt template.

[0085] The image 21 is input to the image encoder 22, where the image 21 can be a sample image in the adjustment process of the candidate probability distribution embedding vector. The feature matching module 30 is connected with the probability prompt learning module 13 and the image encoder 22 respectively, to obtain a matching degree between the probability text embedding vector and the image feature, and then obtain an output 40 based on the probability text embedding vector with the largest matching degree. At this time, the output (which can also be referred to as the result Results) 40 can be a target scene corresponding to the image 21. Alternatively, the feature matching module 30 obtains a sample matching degree between the candidate probability text embedding vector and the sample image feature, and obtains scene prediction information corresponding to the sample image based on the candidate probability text embedding vector with the largest sample matching degree. At this time, the output 40 can be the scene prediction information corresponding to the sample image.

[0086] The text encoder and the image encoder in FIG. 2 are pre-trained and frozen. The text encoder and the image encoder are aligned, for example, through a CLIP (Contrastive Language-Image Pre-training, a pre-training model based on contrastive text-image pairs) or BLIP-2 (an efficient visual-language pre-training method).

[0087] FIG. 3 shows a schematic diagram of the probability prompt learning module 13 in FIG. 2. As shown in FIG. 3, the probability prompt learning module 13 includes a deterministic embedding between class and prompts module 131 and a probabilistic embedding based on Gaussian distribution module 132. The input end of the deterministic embedding between class and prompts module 131 is connected to the output end of the text encoder 12, and can also be decoupled from the text encoder 12. The output end of the deterministic embedding between class and prompts module 131 is connected to the input end of the probabilistic embedding based on Gaussian distribution module 132.

[0088] In an example embodiment, the probability prompt learning module 13 can further include a resampling analysis module 133. The input end of the resampling analysis module 133 is connected to the output end of the probabilistic embedding based on Gaussian distribution module 132. The output end of the resampling analysis module 133 or the output end of the probabilistic embedding based on Gaussian distribution module 132 serves as the output end of the probability prompt learning module 13.

[0089] FIG. 4 is a point site governance system structure based on large model cross entropy Gaussian distribution probability prompt learning for a vision network according to an embodiment of the present disclosure. As shown in FIG. 4, the feature matching module 30 can further include a cross entropy calculation module 31 and a similarity calculation module 32. When the difference (e.g., loss) between the output 40 and the scene label information does not meet a preset condition, it will return to the deterministic embedding between class and prompts module 131 of the probability prompt learning module 13 to adjust the value of the C candidate probability text embedding vectors, for example, to reselect new C prompt templates.

[0090] The method and system provided by the embodiments of the present disclosure will be illustrated below in conjunction with FIGS. 3 and 4.

[0091] As shown in FIG. 3, it is assumed that the prompt template 11 includes K prompt templates. For example, it is assumed that the first prompt template p1 is “this is a picture of a kitchen scene”, the second prompt template p2 is “this is a picture of a street scene”, the third prompt template p3 is “this is a picture of a garbage room scene”, the fourth prompt template p4 is “this is a picture of an elevator scene”, and the Kth prompt template pK is “this is a picture of a construction site scene”.

[0092] In the embodiments of the present disclosure, a plurality of text representations can be used to construct the prompt template P, which is obtained through model learning, and constitutes the prompt templates of K categories in the scene target category or scene category to be judged as shown below:

[0093] Wherein p 1 ,p 2 ,……p K , etc. can represent the following, respectively:

[0094] L in the above formula is the length of the prompt word (i.e. the length of the prompt template), and L is a positive integer greater than or equal to 1. It can be understood that the number of words contained in different prompt templates can be different, but for the prompt template, the scene prompt information in the prompt template such as (kitchen) and (garbage chamber) can be boxed with parentheses, so that the length L seen by the model is the same.

[0095] Here, the image is represented as g, and the image encoder is The text encoder is In this model structure, the image encoder and the text encoder K prompt templates are input into the text encoder 12 to obtain K text vectors. The image 21 is input into the image encoder 22. As shown in FIG. 4, it is assumed here that the input picture of the image encoder 22 (i.e. the image 21) is a garbage chamber monitoring scene. The text encoder For example, it can be a transformer model. The image encoder can use any image encoding method.

[0096] In the embodiments of the present disclosure, when the probability prompt learning module 13 adopts a Gaussian distribution, it can be called a Gaussian distribution probability prompt learning module. The present disclosure designs a Gaussian distribution probability prompt learning module based on a large model. FIG. 3 shows the structure of the Gaussian distribution probability prompt learning module based on a large model designed in the embodiments of the present disclosure. Through the Gaussian distribution probability prompt learning module based on a large model, Gaussian distribution probability prompt learning is used to effectively represent the probability text embedding or probability text embedding vector.

[0097] After the K categories of prompt templates pass through the text encoder, they will pass through the calculation of the deterministic embedding module 31 between the category and the prompt. Subsequently, the probability embedding module 132 based on Gaussian distribution (or called probability embedding module based on factor Gaussian distribution), resampling analysis module 133 are input.

[0098] In the certainty embedding module 131 between the category and the prompt in the embodiments of the present disclosure, 15 prompt templates are randomly selected from 200 prompt templates. C represents the number of scene target categories to be judged, that is, the number of prompt templates selected at a time when feature matching is performed with the image 21. K refers to the number of preset prompt templates in the database, which is usually large (for example, 200) and is frequently updated, and new prompt templates are constantly added. C is initially preset, for example, C = 10, 15, or 20. C is less than K, and the probability prompt learning module 13 selects C from K prompt templates to match with the image 21, because the image 21 usually does not match with K prompt templates, but only matches with a small part of the prompt templates. If 200 prompt templates are directly matched every time, the data calculation amount is large. In the embodiments of the present disclosure, C (for example, 15) prompt templates are randomly selected, and the difference (for example, distance) between the 15 prompt templates and the image 21 is calculated respectively. If the distance between the 15 prompt templates and the image 21 is large, new 15 prompt templates are selected from the 200 prompt templates. The value of C can be adjusted according to fine tuning, that is, training.

[0099] The certainty embedding module 131 between the category and the prompt can be based on a text encoder The output K category prompt templates generate candidate probability text embedding vectors (also referred to as text embedding representations) w for each scene target category c to be judged. c w is represented as a set of prompt template representations of category c: c w is represented as a set of prompt template representations of category c:

[0100] w1, w2, w3, … w c may be represented as follows, respectively:

[0101] The set w c is used to define a probability distribution from which the probability prompt of the scene category c is sampled. For example, initially, w1 = {1, 0, 0, …} is set, that is, the first prompt template in the Kth prompt template. The probability sampling is performed later to make w c change, for example, to w1 = {0.7, 0.001, 0, …}, that is, the values of some dimensions are close to 0. In this way, on the one hand, the extraction of the prompt template becomes regular, and on the other hand, each w c does not simply represent a certain prompt template, but the K prompt templates can be represented by the C prompt templates, but only C vectors need to be calculated when calculating, thereby reducing the data calculation amount.

[0102] In the certainty embedding module 131 between the category and the prompt, each wc represents a K-dimensional row vector. The probability embedding module 132 based on Gaussian distribution processes each row vector in w1, w2, …, to wc, i.e., each row vector is processed as follows. The following takes the cth row vector as an example. C

[0103] In the embodiments of the present disclosure, the probability density function p(z c |w c ) is defined as a factorized Gaussian probability distribution with a center vector and a diagonal covariance matrix (obtained by solving the row vector w c ). The center vector is set as:

[0104] That is, the center vector of the vector w c is itself. The calculation of the factorized Gaussian probability density function p(z c |w c ) is as follows, so that it satisfies the Gaussian distribution, is subject to the factorized Gaussian probability density function, and models the distribution of the prompt learning to be classified as a Gaussian probability model:

[0105] The Gaussian probability model can be understood as the probability distribution of the possible representation to be classified. z c represents the vector obtained after probability sampling of w c using the Gaussian distribution, i.e., the cth candidate probability text embedding vector.

[0106] From p(z c |w c ), C candidate probability text embedding vectors are obtained:

[0107] where z1, z2, z3, …, z C are as follows:

[0108] Then, based on the resampling analysis module 133, z c is re-parameterized, and the purpose of resampling is to make it satisfy the normal distribution and be smoother, and the calculation formula is as follows, i.e., the resampling analysis module 133 extracts z c according to the Gaussian probability part:

[0109] μ(w c ) is the mean of p(z c |w c ), and σ(w c ) is the covariance of p(z​c |w c ) of the standard deviation, and ∈K follows a normal distribution:

[0110] Then it is sent to the subsequent feature matching module 30 (which includes a similarity calculation module 32). As shown in FIG. 4, the embodiment of the present disclosure designs a similarity calculation (Similarity Calculation) module 32 based on the cross entropy of Gaussian distribution probability prompt learning of large models.

[0111] After the resampling analysis module 133, the following is obtained The cross entropy calculation (Cross Entropy Calculation) module 31 calculates the cross entropy between the output features (such as image features or sample image features) of the image encoder and z c , and then the similarity function is calculated through the similarity calculation module 32 to obtain the final scene result to be judged (output 40, i.e. Results), which can be a target scene or scene prediction information.

[0112] When the picture / image 21 is input into the image encoder 22, the feature g d of this picture is obtained d , and the cross entropy is calculated between the feature g c of this picture and each text feature z d . There are a total of C text features, i.e. the value of c is from 1 to C, and the representation of g c is as follows:

[0113] In order to solve formula (9), the vector lengths of z d and g d are kept consistent, and the vector length of g d can be set by the image encoder.

[0114] The θ (d,j) between the feature g C of the picture and each text feature z1, z2, z3,…z (d,j) is calculated, as shown in formula (1) above.

[0115] In formula (1) above, is an element in the vector of the picture feature g d . is an element in the vector of the text feature z c . The q between the picture feature g d and the text feature z c can be calculated.(d,c) The value, i.e., the value of Cross Entropy.

[0116] After obtaining the value of Cross Entropy, the category distance value is calculated, and the smaller the value of Cross Entropy, the more similar the vector representing the text feature and the vector representing the picture feature, and the probability of the image g belonging to the category c in the text z is calculated as follows:

[0117] The distance value of the image belonging to the scene category c to be understood. The category with the smallest distance value, i.e., the scene category to which the image belongs, is the category to be judged. (d,c) The smallest category, i.e., the scene category to which the image belongs, is the category to be judged. The calculated The probability value of the image g belonging to the text category c, and the point site scene information or point site information or scene information of the image collected by the output video networking camera.

[0118] The embodiments of the present disclosure provide a point site management method and system based on cross entropy Gaussian distribution prompt learning of a large model, which uses Gaussian distribution probability prompt learning to fine-tune the parameters of the large model, and introduces a new pixel-text matching loss calculation method to improve the judgment accuracy of the scene category in the video networking. This probability prompt learning method can make the model better understand and process the relationship between vision and language, and has high performance and generalization ability.

[0119] In order to realize the automatic acquisition of the point site information of each camera in the video networking, the embodiments of the present disclosure design a Gaussian probability prompt learning method based on a large model, and design a pixel-text matching loss calculation method based on cross entropy, to realize the intelligent understanding, automatic acquisition and automatic labeling of the point site information of each camera in the video networking. Classify the scene categories of millions of cameras, and push the corresponding resources to the cameras belonging to the same scene category. This system solves the point site management problem of millions of cameras in the video networking. That is, the embodiments of the present disclosure provide a point site management method and system based on cross entropy Gaussian distribution prompt learning of a large model for video networking. It can improve the data processing efficiency in the video networking.

[0120] For the requirements and functional architecture technical route of the general large model generation system in intelligent video monitoring, the problem to be solved is the point site management of millions of cameras in the video networking. Through the Bray Curtis Distance (Bray Curtis Distance) Gaussian distribution probability prompt learning method based on a large model, the point site information of each camera in the video networking is automatically acquired.

[0121] The disclosure belongs to the field of artificial intelligence, and is subdivided into general large models, prompt learning, probabilistic learning, and video networking fields. The disclosure relates to key technology research and application of a multi-modal large model for a video networking operation and maintenance field.

[0122] Based on the same inventive concept, the embodiments of the disclosure also provide an information processing device, as described in the following embodiments. Since the principles of the embodiments of the information processing device solve problems similar to the above-mentioned method embodiments, the implementation of the embodiments of the information processing device can be referred to the implementation of the above-mentioned method embodiments, and the repeated parts will not be described here.

[0123] FIG. 5 shows a structural block diagram of an information processing device in an embodiment of the disclosure. As shown in FIG. 5, the information processing device 500 includes a processing unit 510.

[0124] The processing unit 510 is configured to obtain a probabilistic text embedding vector.

[0125] The processing unit 510 is further configured to obtain an image feature of an image, and obtain a matching degree between the probabilistic text embedding vector and the image feature.

[0126] The processing unit 510 is further configured to take the probabilistic text embedding vector with the largest matching degree as a target probabilistic text embedding vector.

[0127] The processing unit 510 is further configured to determine a target prompt template corresponding to the target probabilistic text embedding vector, and determine a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template.

[0128] Other contents of the embodiment of FIG. 5 can be referred to the above-mentioned embodiments, which will not be described here.

[0129] Based on the same inventive concept, the embodiments of the disclosure also provide an electronic device, as described in the following embodiments. Since the principles of the embodiments of the electronic device solve problems similar to the above-mentioned method embodiments, the implementation of the embodiments of the electronic device can be referred to the implementation of the above-mentioned method embodiments, and the repeated parts will not be described here.

[0130] It should be noted that the above-mentioned modules / units as part of the device can be executed in a computer system such as a group of computer executable instructions.

[0131] Those skilled in the art can understand that each aspect of the disclosure can be implemented as a system, a method or a program product. Therefore, each aspect of the disclosure can be embodied as a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module" or "system" here.

[0132] The electronic device 600 according to this implementation of the present disclosure is described below with reference to FIG. 6. FIG. 6 shows the electronic device 600 as only one example, and should not be taken as limiting the functionality or use of implementations of the present disclosure.

[0133] The electronic device 600 can be a terminal, also referred to as a terminal device, user equipment (UE), mobile station, mobile terminal, etc. The terminal can be widely applied in various scenarios, such as device-to-device (D2D), vehicle to everything (V2X) communication, machine-type communication (MTC), internet of things (IOT), virtual reality, augmented reality, industrial control, autonomous driving, remote medical treatment, smart grid, smart furniture, smart office, smart wear, smart transportation, smart city, etc. The terminal can be a mobile phone, tablet computer, computer with wireless transceiver function, wearable device, vehicle, unmanned aerial vehicle, helicopter, airplane, ship, robot, mechanical arm, smart home device, etc. Embodiments of the present disclosure do not limit the specific technology and specific device form adopted by the terminal.

[0134] As shown in FIG. 6, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 can include, but are not limited to, at least one processing unit 610, at least one storage unit 620, and a bus 630 connecting different system components, including the storage unit 620 and the processing unit 610.

[0135] The storage unit 620 stores program codes that can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary implementations of the present disclosure described in the “Exemplary Methods” section of the present specification.

[0136] The storage unit 620 can include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 6201 and / or a cache memory unit 6202, and can further include a read-only memory (ROM) 6203.

[0137] The storage unit 620 can further include program / utility 6204 having a set of (at least one) program modules 6205, including but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or a combination thereof can include implementation of a network environment.

[0138] Bus 630 can be one of several types of bus structure, including a storage bus or bus for a storage controller, a peripheral bus, a graphics acceleration port, a processor bus, or a local bus using any of a variety of bus architectures.

[0139] Electronic device 600 can also communicate with one or more external devices 640 such as a keyboard or pointing device, a Bluetooth device, etc.; other devices that enable a user to interact with electronic device 600; and / or any devices (e.g., a router, a modem, a network card, etc.) that enable electronic device 600 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface 650. Still yet, electronic device 600 can communicate with one or more networks, such as one or more local area networks (LANs), one or more wide area networks (WANs), and / or the Internet, through network adapter 660. As depicted, network adapter 660 communicates with the other components of electronic device 600 via bus 630. It should be appreciated that although not shown, other hardware and / or software components could be used in conjunction with electronic device 600. These include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0140] Those skilled in the art will readily understand that the example embodiments described herein can be implemented by software and / or by software in combination with the necessary hardware. Thus, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, etc.) or a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to perform the methods according to the embodiments of the present disclosure.

[0141] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer program product, which includes a computer program that, when executed by a processor, implements the information processing method described above.

[0142] In the example embodiments of the present disclosure, a computer readable storage medium is also provided, which can be a readable signal medium or a readable storage medium. The computer readable storage medium has stored thereon a program product capable of implementing the method of the present disclosure. In some possible implementations, various aspects of the present disclosure can also be implemented as a program product in the form of a computer readable medium, which includes program code to cause the terminal / network device to perform the steps described in the above "Example Method" section according to various example embodiments of the present disclosure when the program product is executed on the terminal / network device.

[0143] More specific examples of the computer readable storage medium in the present disclosure can include but are not limited to: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0144] In the present disclosure, the computer readable storage medium can include a data signal carried in the baseband or as part of a carrier wave propagating through the transmission medium, in which a readable program code is borne. Such a propagating data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate or transmit programs for use by or in connection with an instruction execution system, apparatus or device.

[0145] Optionally, the program code contained in the computer readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0146] In specific implementation, the program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., and a conventional procedural programming language such as "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, through the Internet by connecting to an Internet service provider).

[0147] In the present disclosure, at least one (item) can be described as one (item) or multiple (items), and the multiple (items) can be two (items), three (items), four (items), or more (items) without limitation. " / " can represent that the objects before and after the " / " are in an "or" relationship, for example, A / B can represent A or B; "and / or" can be used to describe the existence of three relationships of associated objects, for example, A and / or B, which can represent three cases of A existing alone, A and B existing together, and B existing alone, where A and B can be singular or plural. In order to facilitate the description of the technical solutions of the present application, "first", "second", "A", or "B" and the like can be used to distinguish functionally identical or similar technical features. The "first", "second", "A", or "B" and the like do not limit the quantity and execution order. Moreover, the "first", "second", "A", or "B" and the like do not necessarily mean different. The words "exemplary" or "for example" are used to mean example, illustration, or description, and any design solution described as "exemplary" or "for example" should not be interpreted as more preferred or more advantageous than other design solutions. The words "exemplary" or "for example" are intended to present the relevant concept in a specific manner for understanding.

[0148] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units embodied.

[0149] In addition, although the steps of the method in the present disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. In addition or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be divided into multiple steps, etc.

[0150] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or a network, and includes a number of instructions to make a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) execute the methods according to the embodiments of the present disclosure.

[0151] Other embodiments of the present disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure following the general principles thereof and including such departures from the present disclosure that come within known use or custom in the art to which the present disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

Claims

1. An information processing method characterized by comprising: The method comprises: obtaining a probability text embedding vector; obtaining an image feature of an image, and obtaining a matching degree between the probability text embedding vector and the image feature; taking the probability text embedding vector with the largest matching degree as a target probability text embedding vector; determining a target prompt template corresponding to the target probability text embedding vector, and determining a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template.

2. The method of claim 1, wherein, The probability distribution embedding vector comprises a first number of probability text embedding vectors, and has a second number of dimensions, each dimension indicating a sampling probability of a corresponding prompt template; each prompt template comprises corresponding scene prompt information.

3. The method of claim 2, wherein, The method comprises: obtaining a first number of candidate probability text embedding vectors, wherein the candidate probability distribution embedding vector has a second number of dimensions, each dimension indicating a candidate sampling probability of a corresponding prompt template; obtaining a sample image feature of a sample image and scene label information of the sample image; obtaining a sample matching degree between each candidate probability text embedding vector and the sample image feature respectively; taking the candidate probability text embedding vector with the largest sample matching degree as a target candidate probability text embedding vector; determining a prompt template corresponding to the target candidate probability text embedding vector and scene prompt information corresponding to the prompt template, to determine scene prediction information corresponding to the sample image; obtaining a difference degree between the scene prediction information of the sample image and the scene label information; adjusting the first number of candidate probability text embedding vectors according to the difference degree, to obtain the first number of probability text embedding vectors.

4. The method of claim 3, wherein, The method comprises: obtaining a first number of initial probability text embedding vectors, wherein the initial probability distribution embedding vector has a second number of dimensions, each dimension indicating an initial probability of a corresponding prompt template; taking the initial probability text embedding vector as a center vector of the corresponding candidate probability text embedding vector; obtaining a diagonal covariance matrix of the initial probability text embedding vector; determining a probability density function of the candidate probability text embedding vector based on the initial probability text embedding vector as a Gaussian probability distribution with the corresponding center vector and the diagonal covariance matrix, to obtain the corresponding candidate probability text embedding vector.

5. The method of claim 4, wherein, The method further comprises: resampling each candidate probability text embedding vector to satisfy a normal distribution.

6. The method of claim 3, wherein, The method comprises: inputting the sample image into a frozen and pre-trained image encoder to obtain the sample image feature.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: obtaining a second number of prompt templates; inputting the second number of prompt templates into a frozen and pre-trained text encoder to obtain a second number of text vectors.

8. The method of claim 1, wherein, The image feature comprises K dimensions, K being a positive integer greater than or equal to 1; wherein a cross-entropy between the probability text embedding vector and the image feature is obtained based on a formula as follows, comprising: where c is a positive integer greater than or equal to 1 and less than or equal to C, C is a number of the number of probability text embeddings, C is a positive integer greater than or equal to 1; q (d,c) represents a cross-entropy between the cth probability text embedding vector and the image feature; a kth dimension representing the image feature, k being a positive integer greater than or equal to 1 and less than or equal to K; the k-th dimension in the c-th probability text embedding vector is represented as; the matching degree and the cross entropy are in an inverse correlation relationship.

9. The method of claim 1, wherein, The method further comprises: obtaining the image from an image acquisition device; determining point information of the image acquisition device according to the target scene corresponding to the image.

10. An information processing apparatus, characterized by comprising: The method comprises: a processing unit configured to obtain a probability text embedding vector; The processing unit is further configured to obtain an image feature of the image, and obtain a matching degree between the probability text embedding vector and the image feature. The processing unit is further configured to select a probability text embedding vector with the largest matching degree as a target probability text embedding vector. The processing unit is further configured to determine a target prompt template corresponding to the target probability text embedding vector, and determine a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template.

11. An electronic device, comprising: The processing unit is further configured to obtain an image feature of the image, and obtain a matching degree between the probability text embedding vector and the image feature. The processing unit is further configured to select a probability text embedding vector with the largest matching degree as a target probability text embedding vector. The processing unit is further configured to determine a target prompt template corresponding to the target probability text embedding vector, and determine a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template. The processing unit is further configured to obtain an image feature of the image, and obtain a matching degree between the probability text embedding vector and the image feature. The processing unit is further configured to select a probability text embedding vector with the largest matching degree as a target probability text embedding vector.

12. A computer readable storage medium having stored thereon a computer program, characterized in that, The processing unit is further configured to determine a target prompt template corresponding to the target probability text embedding vector, and determine a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template. The processing unit is further configured to obtain an image feature of the image, and obtain a matching degree between the probability text embedding vector and the image feature. The processing unit is further configured to select a probability text embedding vector with the largest matching degree as a target probability text embedding vector. The processing unit is further configured to determine a target prompt template corresponding to the target probability text embedding vector, and determine a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template. The processing unit is further configured to obtain an image feature of the image, and obtain a matching degree between the probability text embedding vector and the image feature. The processing unit is further configured to select a probability text embedding vector with the largest matching degree as a target probability text embedding vector. The processing unit is further configured to determine a target prompt template corresponding to the target probability text embedding vector, and determine a target scene corresponding to the image according to target scene prompt information corresponding to the target prompt template. The processing unit is further configured to obtain an image feature of