Information retrieval method
By calculating the similarity between the text feature vector of the query information and the image feature vector of the database image frame, determining the target video segment and generating search results, the problem of low retrieval accuracy caused by the limited number of preset event tags is solved, and cross-modal matching and efficient retrieval of text and images are achieved.
Patent Information
- Application Number
- CN202510422686.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The number of preset event tags is limited, which makes it difficult for some video information in the database to be marked. In the actual search process, the query text that is not covered by the preset event tag cannot match the corresponding video information, resulting in low accuracy of the search results.
By obtaining the text feature vector of query information and the similarity value of the image feature vector of each image frame in the database, it is determined whether there is a target video segment in the database. The target video segment includes the target image frame. The similarity value of the text feature vector and the image feature vector is greater than the preset similarity value, and the search results are generated based on the video information of the target video segment.
Cross-modal matching between text and images is realized, effectively solving the missed detection problem caused by the limited number of preset event tags, and improving the accuracy of the search results.
Smart Images

Figure CN119938980A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of information processing technology, and in particular, relates to an information retrieval method. Background Art
[0002] In the application scenario of household robots, household robots usually use their own visual sensors to collect video information and match the video information with preset event tags. When the preset event tags include event tags that match the video information, the video information is annotated with the matching event tags, and the video information and the corresponding event tags are stored in the database together; when the preset event tags do not include event tags that match the video information, the video information is not annotated and is stored in the database. In this way, after receiving the query text input by the user, the household robot will retrieve the event tags that match the query text and the video information corresponding to the event tags in the database, and generate the response content to the user's query based on the event tags and the video information corresponding to the event tags.
[0003] However, the number of event tags included in the preset event tags is relatively limited. These limited event tags cannot comprehensively cover various events in home scenes. In the actual retrieval process, once the database includes video information related to the query information, and the video information is not covered by the preset event tags, the video information matching the query information cannot be found in the database, resulting in missed detection, which in turn leads to low accuracy of the retrieval results. Summary of the invention
[0004] The embodiments of the present application provide an information retrieval method, apparatus, device, medium and product to solve the problem that the limited number of preset event tags makes it difficult to label some video information in the database, resulting in that in the actual retrieval process, the query text not covered by the preset event tags cannot be matched to the corresponding video information, resulting in missed detection and causing the problem of low accuracy of the retrieval results.
[0005] In a first aspect, an embodiment of the present application provides an information retrieval method, which is applied to a cloud server. The information retrieval method includes: Obtain inquiry information; Determine whether there is a target video segment in the database according to a first similarity value between a first text feature vector of the query information and an image feature vector of each image frame in the database, the target video segment includes a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value; When there is a target video segment in the database, a retrieval result for the query information is generated according to the video information of the target video segment.
[0006] In a second aspect, an embodiment of the present application provides an information retrieval method, which is applied to an electronic device. The information retrieval method includes: Receive inquiry information input by the user; Sending a search request to the cloud server, the search request carrying query information, and the search request is used to request the cloud server to search for search results for the query information in the database; receiving a search result sent by a cloud server, where the search result is generated by the cloud server based on video information of a target video segment, where the target video segment is determined by the cloud server based on a first similarity value between a first text feature vector of the query information and an image feature vector of each image frame in a database, where the target video segment includes a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value; Display search results.
[0007] In a third aspect, an embodiment of the present application provides an information retrieval device, which is applied to a cloud server. The information retrieval device includes: A first acquisition module, used to acquire inquiry information; A first determination module is used to determine whether there is a target video segment in the database according to a first similarity value between a first text feature vector of the query information and an image feature vector of each image frame in the database, the target video segment includes a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value; The first generating module is used to generate a retrieval result for the query information according to the video information of the target video segment when there is a target video segment in the database.
[0008] In a fourth aspect, an embodiment of the present application provides an information retrieval device, which is applied to an electronic device, and the information retrieval device includes: A first receiving module, used to receive inquiry information input by a user; A first sending module is used to send a search request to the cloud server, the search request carries query information, and the search request is used to request the cloud server to search for a search result for the query information in the database; a second receiving module, configured to receive a retrieval result sent by the cloud server, the retrieval result being generated by the cloud server based on video information of a target video segment, the target video segment being determined by the cloud server based on a first text feature vector of the query information and a first similarity value of an image feature vector of each image frame in a database, the target video segment including a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame being greater than a first preset similarity value; Display module, used to display search results.
[0009] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, an information retrieval method as described in any one of the first aspects is implemented.
[0010] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, an information retrieval method as described in any one of the first aspects is implemented.
[0011] In a seventh aspect, an embodiment of the present application provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, an information retrieval method as described in any one of the first aspects is implemented.
[0012] The information retrieval method, device, equipment, medium and product of the embodiment of the present application can perform similarity calculation based on the first text feature vector of the query information and the image feature vector of each image frame in the database, and the similarity value between the first text feature vector and the image feature vector of each image frame reflects the degree of association between the semantic connotation contained in the query information and the visual content presented by the image frame. Based on the first preset similarity value, the target video segment corresponding to the target image frame matching the query information is retrieved from the database, and the cross-modal matching of text and image is realized. Then, according to the video data corresponding to the target video segment, the retrieval result for the query information is generated. In this way, through this cross-modal matching method based on the similarity calculation of the text feature vector and the image feature vector, the problem that some video information in the database is difficult to be labeled due to the limited number of preset event tags is effectively solved, and then in the actual retrieval process, the query text not covered by the preset event tag cannot be matched to the corresponding video information, and finally misses detection and causes the problem of low accuracy of the retrieval result. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0014] Figure 1 A schematic diagram showing a process of applying the information retrieval method provided in some embodiments of the present application to a cloud server; Figure 2 A flowchart showing a specific implementation of step 120 provided in some embodiments of the present application is shown; Figure 3A flowchart showing another specific implementation of step 120 provided in some embodiments of the present application is shown; Figure 4 A flowchart diagram showing another specific implementation of step 120 provided in some embodiments of the present application is shown; Figure 5 A flowchart showing a specific implementation of step 420 provided in some embodiments of the present application is shown; Figure 6 A schematic flow chart of a method for building a database in an information retrieval method provided in some embodiments of the present application is shown; Figure 7 A schematic diagram showing a flow chart of an information retrieval method provided by some embodiments of the present application applied to an electronic device; Figure 8 A schematic flow chart of a method for sending reference video information to a cloud server in an information retrieval method provided in some embodiments of the present application is shown; Fig. 9 A schematic diagram showing the structure of an information retrieval device provided in some embodiments of the present application applied to a cloud server; Fig.10 A schematic diagram showing the structure of an information retrieval device provided by some embodiments of the present application applied to an electronic device; Fig.11 A schematic diagram of the structure of an electronic device provided in some embodiments of the present application is shown. DETAILED DESCRIPTION
[0015] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0016] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0017] It should be noted that the acquisition, storage, use and processing of data in the embodiments of the present application are in compliance with the relevant provisions of national laws and regulations.
[0018] It should be noted that in the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned, and they should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.
[0019] In order to solve the problems in the above-mentioned related technologies, the embodiments of the present application provide an information retrieval method, device, equipment, medium and product. Figure 1 To Attachment Figure 8 , the information retrieval method provided by the embodiment of the present application is described in detail through specific embodiments and their application scenarios.
[0020] Figure 1 FIG. 1 is a flow chart showing an information retrieval method provided by some embodiments of the present application. Figure 1 As shown, the information retrieval method may include steps 110 to 130.
[0021] Step 110, obtaining inquiry information.
[0022] Step 120, based on the first similarity value between the first text feature vector of the query information and the image feature vector of each image frame in the database, determine whether there is a target video segment in the database, the target video segment includes a target image frame, and the first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value.
[0023] Step 130: if there is a target video segment in the database, generate a search result for the query information according to the video information of the target video segment.
[0024] Thus, by calculating the similarity between the first text feature vector of the query information and the image feature vector of each image frame in the database, the similarity value between the first text feature vector and the image feature vector of each image frame reflects the degree of association between the semantic connotation contained in the query information and the visual content presented by the image frame. Based on the first preset similarity value, the target video segment corresponding to the target image frame matching the query information is retrieved from the database, and the cross-modal matching of text and image is realized. Then, according to the video data corresponding to the target video segment, the retrieval result for the query information is generated. In this way, through this cross-modal matching method based on the similarity calculation of the text feature vector and the image feature vector, the problem that some video information in the database is difficult to be labeled due to the limited number of preset event tags is effectively solved, and then in the actual retrieval process, the query text not covered by the preset event tag cannot be matched to the corresponding video information, and finally misses the detection and causes the problem of low accuracy of the retrieval result.
[0025] The above steps are described in detail below, as shown below.
[0026] First, step 110 is involved. The query information involved in the embodiment of the present application refers to text information or voice information input by the user through the electronic device for searching and finding relevant content in the database. For example, in an information retrieval system, a sentence such as "What did the first user at home do today" input by the user is the query information.
[0027] For example, the query information can be provided by the user manually typing text content into the input box of the retrieval system; the query information can also be provided by using voice recognition technology, where the user speaks the content he wants to retrieve to a device with voice input function such as a smart phone, smart speaker, smart robot, etc., and the electronic device converts the voice into corresponding text information as the query information.
[0028] Secondly, in step 120, the first text feature vector is a vector representation obtained by processing the text through a specific mathematical method. The first text feature vector can reflect the semantics, grammar and other information in the text in the form of a digital vector. For example, for the query information "What did the first user at home do today?", after the corresponding feature extraction algorithm such as the pre-trained Contrastive Language - Image Pretraining (CLIP) model, it is mapped into a text feature vector in the form of latent space encoding, that is, the text content related to "home" and "first user" is reflected in the form of a numerical vector. It can be understood that different text contents correspond to different vectors, and similar text vectors will also be relatively close in space.
[0029] The image feature vector is a vector representation obtained after feature extraction of an image, including feature information such as color, texture, and shape of the image frame. For example, for an image frame of "the first user cleaning at home", the image feature vector corresponding to the image frame is mapped through a corresponding feature extraction algorithm such as a pre-trained CLIP model. This image feature vector integrates various feature information related to the scene of the first user cleaning at home in the image frame, such as color, such as the color and texture of the room, and shape, such as the first user's posture outline, the shape of the cleaning object, etc., in the form of numerical values to form a vector representation that can represent the features of the image frame.
[0030] The first similarity value refers to the similarity value between the first text feature vector corresponding to the query information and the image feature vector of each image frame in the database, which is obtained by using similarity calculation methods such as cosine similarity, Euclidean distance and other algorithms to measure the similarity by calculating the distance or angle between the text feature vector and the image feature vector, and is used to determine the matching situation between each image frame and the query information. For example, the first similarity value between the image feature vector of each image frame and the first text feature vector is obtained by using the CLIP model using the cosine similarity calculation method or the Euclidean distance calculation method.
[0031] The first preset similarity value is a pre-set numerical standard for measuring whether the similarity meets the standard. When the first similarity value of the first text feature vector and the image feature vector is greater than the first preset similarity value, the two are considered to be sufficiently similar, and the corresponding image frame is a target image frame that meets the retrieval requirements.
[0032] The target video segment refers to a video segment including a target image frame. The target video segment is a retrieved video content portion related to the query information. The target image frame is a specific image frame in the target video segment that has a high degree of similarity in features with the query information.
[0033] In some embodiments of the present application, before the above-mentioned step 120, the above-mentioned information retrieval method may further include obtaining an event type corresponding to the query information, the event type includes a first event type or a second event type, the occurrence frequency of the first event type is higher than the occurrence frequency of the second event type, and the danger level of the first event type is lower than the danger level of the second event type, and the event type is determined by at least one of the following methods: determined based on a video tag matched with the query information, and determined based on keywords extracted from the query information.
[0034] Among them, the event type is a classification and summary of the content involved in the inquiry information, which is used to distinguish events of different nature and characteristics. The frequency of occurrence refers to the frequency of a certain type of event in a specific environment or time period. The degree of danger refers to the size of the risk or harm that a certain type of event may bring. It can be understood that the first type of event is a regular event, which is more common and occurs commonly in daily home environments or scenes monitored by home care robots, and usually does not pose a threat or harm to the safety of family members, system stability or other key aspects. The second type of event is a warning event. In contrast to regular events, this type of warning event is those that occur less frequently, but once they occur, they may have a serious impact or harm on the safety of life and property of people, normal operation, etc. For the event corresponding to the inquiry information, if the event is not in the pre-constructed warning list, it will be determined as a regular event. Among them, the pre-constructed warning list is a list that has been sorted, analyzed and determined in advance, and can include various specific event descriptions, keywords, characteristics and other information identified as warning events. Exemplarily, in the application scenario of a home care robot, the pre-built warning list may include but is not limited to the second user falling down, the first user covering his face, the dog leaving the fence, the dog getting on the table, electrical appliances smoking and catching fire, and a stranger breaking in.
[0035] In one example, the acquisition of the event type corresponding to the query information can be determined based on the keywords extracted from the query information, which can specifically include performing text preprocessing operations on the query information, such as word segmentation, part-of-speech tagging, and then identifying the named entities in the query information, such as names of people, animal names, actions or event descriptions. Then, after identifying the named entities, the relationship between the entities is extracted to form a complete event description. Then, the event description is compared with the pre-built alert list to determine whether the event corresponding to the query information is an alert event. If the event is in the alert list, it is classified as an alert event, i.e., the second event type. If the event is not in the alert list, it is classified as a regular event, i.e., the first event type.
[0036] Video tags are a way of marking or classifying the video content of a video segment, which is used to summarize the key information in the video segment. Video tags can include, but are not limited to, event tags, time tags, and video object tags. Based on algorithms such as keyword matching or semantic matching algorithms, the preprocessed query information is matched with the video tags in the database to find out whether there are video tags containing these keywords or highly similar to them in the video tag library of the database. When the query information successfully matches a certain video tag, the event type corresponding to the video tag is determined as the event type corresponding to the query information according to the preset video tag and event type association information. For example, if the matched video tag is "the first user plays", according to the association rule, the event type corresponding to the video tag is determined as the first event type, i.e., a regular event; if the matched video tag is the "kitchen fire" tag, according to the association rule, the event type corresponding to the video tag is determined as the second event type, i.e., a warning event.
[0037] For different event types, based on this, the above step 120 may specifically include determining whether there is a target video segment in the database according to the first similarity value between the event type and the first text feature vector of the query information and the image feature vector of each image frame in the database.
[0038] Exemplarily, after the event type is acquired and the first similarity value between the first text feature vector of the query information and the image feature vector of each image frame in the database is calculated using the CLIP model, the information is combined to determine whether the target video segment exists. Different screening strategies are set for different event types. For example, for the first event type with high frequency of occurrence and low degree of danger, when determining whether the similarity value meets the requirements, the similarity threshold, i.e., the first preset similarity value, can be appropriately relaxed. For the second event type with low frequency of occurrence and high degree of danger, a more accurate match is required, so the similarity threshold, i.e., the first preset similarity value, is increased. Only when the first similarity value is higher than the higher first preset similarity value, the video segment containing the corresponding image frame is determined to be the target video segment.
[0039] Therefore, by first determining the event type corresponding to the query information, and then determining the target video segment in combination with the first preset similarity value corresponding to the event type, the indiscriminate and uniform standard search of the entire database is avoided. Through this differentiated first preset similarity value method, it is possible to better match the target video segment that truly meets the query information and is consistent with the nature of the event. For example, for the second event type, a higher similarity threshold can ensure that the retrieved video segment is indeed highly consistent with the event described in the query information, reducing false matches, thereby improving the accuracy and reliability of the retrieval results, and making the retrieval results more in line with the content that the user actually wants to find.
[0040] In some embodiments of the present application, the event type includes a first event type, the first preset similarity value includes a first reference similarity value, the target image frame includes a first target image frame, and a first similarity value between the first text feature vector and the image feature vector of the first target image frame is greater than the first reference similarity value. Based on this, the above step 120 may specifically: Determine whether there is a first target image frame in the database according to a first text feature vector of the query information and a first similarity value of an image feature vector of each image frame in the database; In the case that there is a first target image frame in the database, it is determined that there is a target video segment in the database; or, in the case that there is no first target image frame in the database, it is determined that there is no target video segment in the database.
[0041] The first reference similarity value is a pre-set measurement standard value corresponding to the first event type, which is used to determine whether the similarity between the first text feature vector and the image feature vector meets the requirements of the first event type, and then determine whether the first target image frame that meets the query information retrieval condition is found. For example, the first reference similarity value can be 65%. The first target image frame is an image frame retrieved from multiple image frames in the database, and the first similarity value between the first text feature vector corresponding to the query information and the image feature vector of the first target image frame is greater than the first reference similarity value.
[0042] Exemplarily, for the query information "Did the second user watch TV today?", the query information is processed to obtain the event "the second user watches TV" corresponding to the query information, and the above-mentioned CLIP model is used to encode it to obtain a first text feature vector, and the first text feature vector is searched for similarity with the image feature vectors of the image frames stored in the database to obtain a retrieval result, that is, a first similarity value between the first text feature vector and the image feature vector of each image frame. If the first reference similarity value is set to 65%, the results whose first similarity value is greater than the first reference similarity value are screened out from the retrieval results to obtain multiple data entries, such as [11:30, second user, second user watches TV, 90%], [10:00, kitchen, second user prepares to watch TV, 70%], and finally, the data entry with the highest confidence is selected, and a retrieval result for the query information is integrated according to the video information therein, such as [11:30, second user, second user watches TV, 90%].
[0043] Therefore, by distinguishing different event types from the first event types and setting corresponding first reference similarity values, the first target image frames and target video segments matching the query information can be screened out more accurately.
[0044] In some embodiments of the present application, the event type includes a second event type, the first preset similarity value includes a second reference similarity value, the target image frame includes a second target image frame, the first similarity value between the first text feature vector and the image feature vector of the second target image frame is greater than the second reference similarity value, and the second reference similarity value is greater than the first reference similarity value. Based on this, if Figure 2 As shown, the above step 120 may specifically include step 210 and step 240.
[0045] Step 210 : Compare the first text feature vector of the query information with the sound feature vector of each audio information in the database to obtain a second similarity value between the first text feature vector and the sound feature vector of each audio information.
[0046] Among them, the sound feature vector refers to a numerical vector that is obtained by converting various attributes of audio information such as audio frequency, amplitude, timbre, rhythm, etc. through an audio processing algorithm.
[0047] Exemplarily, a similarity calculation method such as cosine similarity, Euclidean distance, etc. is used to compare and calculate the first text feature vector of the query information with the sound feature vector of each audio information.
[0048] Step 220 , determining whether there is a second target image frame in the database based on the first text feature vector of the query information and the first similarity value of the image feature vector of each image frame in the database.
[0049] Step 230, determining whether there is target audio information in the database based on the second similarity value between the first text feature vector and the sound feature vector of each audio information, wherein the second similarity value between the first text feature vector and the sound feature vector of the target audio information is greater than the second reference similarity value.
[0050] Step 240: If there is a second target image frame and / or target audio information in the database, determine the video segment corresponding to the second target image frame and / or target audio information as the target video segment.
[0051] Exemplarily, the query information is first converted into a first text feature vector, and then compared with the image feature vector of each image frame in the database to calculate a first similarity value. If the first similarity value corresponding to a certain image frame is greater than the second reference similarity value, the image frame is determined to be a second target image frame. After obtaining the second similarity value between the first text feature vector and the sound feature vector of each audio information, if the second similarity value corresponding to the sound feature vector of a certain audio information is greater than the second reference similarity value, the audio information is determined to be the target audio information. When the second target image frame and / or the target audio information exists in the database, the video segment corresponding to the second target image frame and / or the video segment corresponding to the target audio information is determined as the target video segment. That is, if there are matching second target image frames and target audio information at the same time, the video segment corresponding to the second target image frame and the video segment corresponding to the target audio information can both be determined as the target video segment. Alternatively, only one of the matching situations can be selected according to specific needs, such as determining the video segment corresponding to the second target image frame as the target video segment or determining the video segment corresponding to the target audio information as the target video segment.
[0052] Therefore, for the second event type, by combining the image feature vector and the sound feature vector for comprehensive retrieval, the event information can be captured more comprehensively, and by mutual supplementation and verification of different modal information, the target video segment that meets the requirements can be located more accurately.
[0053] In some embodiments of the present application, the query information includes a location keyword, wherein the location keyword is a word or phrase related to the location where the event occurred extracted from the query information. For example, for the query information "What did the second user do in the kitchen", "kitchen" is the location keyword. The event type includes a second event type, the first preset similarity value includes a second reference similarity value, the target image frame includes a second target image frame, and the first similarity value between the first text feature vector and the image feature vector of the second target image frame is greater than the second reference similarity value. Based on this, if Figure 3 As shown, the above step 120 may specifically include steps 310 to 330.
[0054] Step 310 : Compare the second text feature vector of the location keyword with the third text feature vector of each location tag in the database to obtain a third similarity value between the second text feature vector and the third text feature vector of each location tag.
[0055] Exemplarily, the second text feature vector refers to processing the position keyword through a text feature extraction method such as a word vector model, a text encoding model, etc., so as to convert the semantics, grammar and other information of the position keyword into a numerical vector. For example, the second text feature vector obtained after the position keyword "kitchen" is converted can reflect the numerical features related to the kitchen in the vector space. The third text feature vector is a numerical vector obtained after feature extraction of each position tag in the database using the same algorithm as the second text feature vector. The position tag is an identifier of the location information of the event recorded in the video segment. These position tags are converted into the third text feature vector through text feature extraction technology, so that different position tags have a quantifiable representation in the vector space, so as to be compared and analyzed with the second text feature vector of the position keyword.
[0056] Step 320, based on the third similarity value between the second text feature vector and the third text feature vector of each location tag, determine whether there is a target location tag in the database, the third similarity value between the second text feature vector and the third text feature vector of the target location tag is greater than the second reference similarity value, and the location tag is an identifier of the location information of the event recorded in the video segment corresponding to the location tag.
[0057] Exemplarily, each calculated third similarity value is compared with the second reference similarity value. If the third similarity value corresponding to a certain position tag and the second text feature vector is greater than the second reference similarity value, then the position tag is determined as the target position tag.
[0058] Step 330: When there is a second target image frame and / or a target position tag in the database, determine the video segment corresponding to the second target image frame and / or the target position tag as the target video segment.
[0059] Exemplarily, when there is a second target image frame and / or a target position tag in the database, the video segment corresponding to the second target image frame and / or the video segment corresponding to the target position tag is determined as the target video segment. If there are both a matching second target image frame and a video segment corresponding to the target position tag, both can be used as the target video segment; or, according to specific needs, only one of the matching situations, the video segment corresponding to the second target image frame or the video segment corresponding to the target position tag, can be selected as the target video segment.
[0060] Therefore, when the query information for the second event type includes location keywords related to the location, by achieving accurate location matching, it is possible to effectively prevent retrieval errors caused by misjudgment of the location, and ensure that the retrieved video clips accurately reflect the relevant events occurring at the specified location. Furthermore, by combining the accurate matching of location keywords with location tags, the scope of the search can be further narrowed, thereby significantly improving the accuracy of the search results.
[0061] In some embodiments of the present application, before the above step 120, as Figure 4 As shown, the information retrieval method may further include steps 410 to 430.
[0062] Step 410: Obtain keywords of the query information and reference event tags corresponding to the keywords.
[0063] The reference event tag is a guiding identifier associated with the keywords of the query information and used for preliminary screening of video tags.
[0064] Exemplarily, determining corresponding reference event labels according to extracted keywords may be achieved through a pre-established keyword-reference event label mapping table.
[0065] Step 420, based on the reference event tag, searching the database for video tags corresponding to the reference event tag to determine whether there is a target video tag in the database, wherein the video tag includes a sample event tag and / or a sample event attribute tag.
[0066] Among them, the sample event label is a type of video label in the database, which is used to classify and mark the events recorded in the video, so as to quickly identify the approximate event category of the video segment. For example, for a database of a home care robot, the sample event label may include but is not limited to the second user exercising, the second user reading, the first user sleeping, the first user playing, express parcels, the cat on the sofa, the cat and the first user playing, the dog eating, and the dog playing with the ball. The sample event attribute label is a label that annotates the various attribute features of the events recorded in the video segment. The sample event attribute label may include but is not limited to the time attribute label of the event, such as the morning period and the evening period, the location attribute label, such as the bathroom and the kitchen, and the character attribute label, such as the first user and the second user.
[0067] Step 430: If there is a target video tag in the database, determine the video segment corresponding to the target video tag as the target video segment.
[0068] Therefore, by first obtaining reference event tags based on keywords and performing video tag retrieval, a portion of video segments that meet the requirements can be quickly screened out. Because video tags are a pre-summary and classification of video content, this video tag-based retrieval method is relatively simple and efficient, and can directly locate relevant video resources without the need for complex feature vector similarity calculations, effectively avoiding large-scale feature vector calculations and comparisons of the entire database, thereby saving retrieval time and computing resources and improving retrieval efficiency.
[0069] Based on this, the above step 120 may specifically include, when there is no target video tag in the database, determining whether there is a target video segment in the database according to a first similarity value between the first text feature vector of the query information and the image feature vector of each image frame in the database.
[0070] Therefore, this method of first combining the retrieval of reference event tags and video tags and then searching based on feature vectors can improve the accuracy of the retrieval results while making the information retrieval system more adaptable.
[0071] In one example, in response to the query information "What did the first user at home do today?", a search is performed on sample event tags in the database based on the subject of the event in the query information, such as the first user, and the search results obtained include time and event, such as [9:10, first user, first user exercise, 90%] and [13:30, first user, first user sleep, 82%].
[0072] In another example, in response to the query information "What did the second user do in the living room today?", the location tags in the database are retrieved based on the location information of the event in the query information, such as the living room, and the retrieval results include time and events, such as [10:30, second user, second user exercising, 93%] and [13:30, second user, second user reading, 93%].
[0073] In another example, in response to the query "Has the dog been to the table today?", a search is performed on sample event tags in the database based on events such as "Dog is at the table", and the search result is empty.
[0074] In some embodiments of the present application, Figure 5 As shown, the above step 420 may specifically include step 4201 and step 4202.
[0075] Step 4201, obtain the fourth similarity value of the event type corresponding to the reference event tag and the video tag corresponding to the reference event tag in the database, the fourth similarity value is used to characterize the similarity between the video tag and the video segment corresponding to the video tag, the event type includes the first event type or the second event type, the occurrence frequency of the first event type is higher than the occurrence frequency of the second event type, and the danger level of the first event type is lower than the danger level of the second event type.
[0076] The fourth similarity value refers to a value used to measure the similarity between a video tag and a corresponding video segment, and reflects the similarity between the information conveyed by the video tag and the actual content presented by the video segment.
[0077] Exemplarily, a mapping relationship table between reference event labels and event types is pre-established, and the table is queried to determine whether the event type corresponding to the current reference event label is the first event type or the second event type. For example, the reference event labels corresponding to the first event type include but are not limited to the second user exercising, the second user reading, the first user sleeping, the first user playing, express parcels, the cat on the sofa, the cat and the first user playing, the dog eating, and the dog playing with a ball; the reference event labels corresponding to the second event type include but are not limited to the second user falling down, the first user covering his face, the dog out of the fence, and the dog on the table.
[0078] Step 4202, based on the reference event tag, event type and the fourth similarity value of the video tag corresponding to the reference event tag in the database, retrieve the video tag corresponding to the reference event tag in the database to determine whether there is a target video tag in the database, and the fourth similarity value corresponding to the target video tag is greater than the second preset similarity value, and the second preset similarity value is determined based on the event type.
[0079] The second preset similarity value includes a third reference similarity value corresponding to the first event type or a fourth reference similarity value corresponding to the second event type.
[0080] For example, if it is the first event type, since it has a high frequency of occurrence and a low degree of danger, a relatively loose screening condition can be set; if it is the second event type, a stricter screening condition is set, that is, the third reference similarity value is less than the fourth reference similarity value. For example, for the first event type, the third reference similarity value can be set at a relatively low reference similarity value such as 80%, and for the second event type, the fourth reference similarity value can be set at a relatively high reference similarity value such as 90%.
[0081] It should be noted that the third reference similarity value may be the same as or different from the first reference similarity value. Similarly, the fourth reference similarity value may be the same as or different from the second reference similarity value, which is not specifically limited here.
[0082] Therefore, by introducing event type and the fourth similarity value for video tag retrieval, it is possible to more accurately screen out target video tags related to the query information, and in the process of video tag retrieval, by setting different screening strategies for different event types, such as the second preset similarity value, the retrieval process is made more targeted. For example, for the more conventional first event type, the more relaxed third reference similarity value is used to quickly locate possible related video tags, thereby improving retrieval efficiency; and for the more important second event type, the more stringent fourth reference similarity value is used to screen out highly accurate video tags, thereby reducing false matches and improving the accuracy of the retrieval results.
[0083] In some embodiments of the present application, Figure 6 As shown, the above information retrieval method may further include steps 510 to 540.
[0084] Step 510, receiving reference video information sent by an electronic device, the reference video information includes video segments, video identification tags, audio information, image frames and sample event attribute tags, the sample event attribute tags include time tags, location tags, video object tags and warning audio tags.
[0085] The video identification tag is a mark or number used to uniquely identify a video segment. Specifically, it can be a string of numbers and letters, which contains date information and randomly generated characters to ensure its uniqueness. The sample event tag is generated based on the sample event attribute tag, and is a tag that generally describes the overall nature and category of the event involved in the video segment. The audio information includes the audio segment corresponding to the video segment. The video segment is a dynamic visual content formed by a series of continuous image frames according to a certain time, and the image frame is the key image frame that constitutes the video segment. The sample event attribute tag is a tag that annotates various attribute features of the event recorded by the video. The time tag is to clearly record the time-related information of the event in the video, such as the specific year, month, day, hour, minute, and second, or a certain time period range, which is used to locate and distinguish the events in the video segment from the time dimension. The location tag is used to indicate the geographical location or specific place where the event in the video segment occurs. The video object tag is used to annotate objects such as people and objects that appear in the video, such as "first user", "puppy", "kitten", etc. The warning audio tag is used to identify whether there is a warning sound in the video and the type of the warning sound, such as "fire alarm sound", "theft alarm sound", etc.
[0086] Step 520: Generate a sample event tag for the video segment according to the sample event attribute tag.
[0087] Exemplarily, semantic analysis is performed on each attribute information in the sample event attribute label, and then the sample event label is generated based on the semantic analysis result of the sample event attribute label according to the preset label generation rule. Among them, the label generation rule can be formulated based on artificial experience. In the case where there are sample event attribute labels for "exercise", "reading", "eating", "watching TV" and "pet out of the cage" on the image frame, the preset event label corresponding to the sample event attribute label will be determined when the sample event label is generated.
[0088] Step 530: perform feature extraction and vector conversion on the image frame to generate an image feature vector of the image frame.
[0089] Exemplarily, an image feature extraction model or algorithm is used to extract features from the preprocessed image frame, and the extracted image features are converted into image feature vectors in vector form.
[0090] Step 540 , building a database by associating and storing the reference video information, the sample event labels, and the image feature vectors of the image frames.
[0091] Exemplarily, the video identification tag in the reference video information of each video segment is used as the key, and correspondingly, the video segment, audio information, image frame, time tag, position tag, video object tag, warning audio tag and image feature vector of the image frame related to the video identification tag are associated and stored as values.
[0092] Therefore, by integrating multi-source information, generating sample event labels, extracting image feature vectors, and building an associated database, the efficiency and accuracy of information retrieval can be improved.
[0093] In some embodiments of the present application, a regular cleanup operation can be set for the database. Specifically, the cleanup cycle is determined according to the number of days of data storage set by the user. For example, if the user sets the data storage period to 30 days, the cloud server can check whether there is expired data every other day. For the identified expired data, the delete operation provided by the database can be used to clear it.
[0094] The present application also provides an information retrieval method, which is applied to an electronic device, wherein the electronic device and a cloud server establish a communication connection via Socket communication, such as Figure 7 As shown, the information retrieval method may further include steps 610 to 640.
[0095] Step 610: receiving inquiry information input by the user.
[0096] Exemplarily, the electronic device can provide the user with a display interface for inputting query information. For example, a text input box is configured on the display interface of a home care robot. The user can call out the keyboard by touching the screen and enter the content they want to query in the input box. At the same time, the display interface can also set some prompt information to guide the user to accurately express the query needs. For example, the prompt information can be "Please enter the video-related content you want to find" and so on.
[0097] Step 620: Send a search request to the cloud server. The search request carries query information. The search request is used to request the cloud server to search for search results for the query information in the database.
[0098] Step 630, receiving the retrieval result sent by the cloud server, the retrieval result is generated by the cloud server based on the video information of the target video segment, the target video segment is determined by the cloud server based on the first text feature vector of the query information and the first similarity value of the image feature vector of each image frame in the database, the target video segment includes the target image frame, and the first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than the first preset similarity value.
[0099] Step 640, display the search results.
[0100] For example, if the query information input by the user is "Did the second user watch TV today?", a search operation is performed using the above database based on the keywords in the query information such as "second user" and "watch TV", and search results such as [11:30, second user, second user watches TV, 90%] are obtained, and "11:30, the second user watched TV" is displayed through screen display or voice output.
[0101] Therefore, by using cloud servers to perform database retrieval operations, electronic devices do not need to have strong local computing capabilities and large-capacity storage to process video data and perform complex text and image feature vector calculations. By handing over information retrieval tasks to cloud servers, they can fully utilize the advantageous resources of cloud servers to achieve faster and more accurate retrieval.
[0102] In some embodiments of the present application, the electronic device may include an image acquisition module, an audio acquisition module, a positioning acquisition module, and a time acquisition module. Figure 8 As shown, the above information retrieval method may further include steps 710 to 730.
[0103] Step 710: Acquire initial video information, where the initial video information includes an initial video segment, initial audio information, initial time information, and initial position information.
[0104] Among them, the initial video information is an unprocessed information set obtained by the electronic device, which contains various aspects of original data related to the video. Specifically, it may include the initial video segment collected by the image acquisition module, the initial audio information collected by the audio acquisition module, the initial position information collected by the positioning acquisition module, and the initial time information collected by the time acquisition module.
[0105] Step 720, preprocess the initial video information to obtain reference video information, which includes video segments, video identification tags, audio information, image frames and sample event attribute tags. The sample event attribute tags include time tags, location tags, video object tags and warning audio tags.
[0106] Among them, the data preprocessing of the initial video segment can include extracting image frames from the video segment through an image frame extraction module of an electronic device, so as to reduce the frame rate of the video segment while ensuring that the integrity of the information contained in the video is effectively maintained. Specifically, for static video segments in the home scene monitoring video segment, the motion vector of the image frame is calculated for judgment. Once the motion vector is less than a preset threshold value, the video segment is determined to be a static video segment. At this time, the frame extraction frequency can be appropriately reduced, thereby effectively improving storage efficiency. For example, when the home monitoring screen remains static for a long time, such as a room scene where no one is active, unnecessary frame extraction operations are reduced to avoid wasting storage resources. For dynamic video segments, in order to further optimize the frame extraction process, the information similarity is calculated for screening. If the similarity between a frame and a selected key frame is higher than a specific threshold, it is determined that the frame does not need to be extracted, which can also reduce the frame extraction frequency. For example, in a continuous video of a person walking, when the person's movement speed is slow and the background is relatively fixed, the visual difference between adjacent frames is small. The similarity calculation can reduce the extraction of similar frames, thereby reducing the amount of data storage, improving storage efficiency and reducing the burden of subsequent data processing while ensuring the retention of key dynamic information.
[0107] The time information corresponding to a single image frame is obtained by updating the recording node using the timestamp on the electronic device side. During the processing, the node not only records the time of a single image frame, but also detects the frame extraction frequency and calculates the similarity of the image structure to determine whether to delete the current frame, and calculates the time length of the segment. By processing and integrating these time-related data, a time tag that can accurately describe the time range or specific time point of the video event is finally determined. For example, if a video segment is processed by frame extraction, its start time and end time are determined as the corresponding time period information, and it is set as the time tag corresponding to the video segment.
[0108] The room location mark classification node on the electronic device side and the navigation system of the electronic device are used to obtain spatial navigation location information, and combined with the home map for comprehensive analysis, to accurately determine the spatial location information of the image frame corresponding to the current video segment, such as the specific room names such as the living room and kitchen. These determined room name information constitutes the location label. For example, when the robot captures a video segment in the kitchen area, after being processed by the room location mark classification node, the video segment will be labeled with the location label "kitchen".
[0109] The human body detection model node on the electronic device side is used to detect the image frame of the video segment, and individuals of different categories such as the first user and the second user in the image frame are obtained, and each individual is labeled and indexed, and the key point features of the human body structure are extracted as the feature vector of the individual. These detected character category information and individual feature vectors can be used to construct the character part in the video object label. For example, if a second user and a first user are detected in an image frame, then "the second user with index X" and "the first user with index Y" will be recorded in the video object label. In addition, the object detection model node on the electronic device side is used to process the image frame and the preset word list of the video segment to detect whether the image frame contains objects and quantities in the preset word list. For example, if there are items such as "vase" and "mobile phone" in the preset word list, when a vase is detected in the image frame, the information of "vase (quantity 1)" will be added to the video object label. Therefore, through the collaborative work of the human body detection model node and the object detection model node, the video object information such as people and objects appearing in the video can be fully annotated, thereby forming a complete video object label.
[0110] The audio warning classification node on the electronic device side is used to process the time information provided by the node on the electronic device side and the audio information corresponding to the time period of the video segment corresponding to the image frame. Specifically, the automatic speech recognition (ASR) technology can be used to use the hidden Markov model as the acoustic model to convert the initial sound encoding into classification label information. By comparing and matching with the warning sound list such as "smoke alarm", "cat cry", "dog barking", "crackling sound", "burning sound", etc., if the audio information successfully matches a certain type of warning sound in the list, the corresponding warning audio label will be marked. For example, when it is detected that the audio information matches the sound of "smoke alarm", the video segment will be marked with the warning audio label of "smoke alarm".
[0111] Step 730: Send reference video information to the cloud server.
[0112] Therefore, through data preprocessing operations, the original initial video information is converted into high-quality, uniformly formatted and detailed reference video information, and the reference video information is sent to the cloud server. This enables the preprocessed video information to provide more reliable data support for the retrieval algorithm on the cloud server, reducing erroneous or inaccurate retrieval results caused by data quality issues and improving the retrieval quality of the entire information retrieval system.
[0113] Based on the information retrieval method provided in the above embodiment, the present application also provides a specific implementation of the information retrieval device. Please refer to the following embodiment.
[0114] See also Fig. 9 The information retrieval device 800 provided in the embodiment of the present application is applied to a cloud server, and the information retrieval device 800 includes: A first acquisition module 810 is used to acquire inquiry information; A first determination module 820 is used to determine whether there is a target video segment in the database according to a first similarity value between a first text feature vector of the query information and an image feature vector of each image frame in the database, the target video segment includes a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value; The first generating module 830 is used to generate a search result for the query information according to the video information of the target video segment when there is a target video segment in the database.
[0115] Thus, the first determination module 820 performs similarity calculation on the first text feature vector of the query information obtained by the first acquisition module 810 and the image feature vector of each image frame in the database, and the obtained similarity value between the first text feature vector and the image feature vector of each image frame reflects the degree of association between the semantic connotation contained in the query information and the visual content presented by the image frame. Based on the first preset similarity value, the target video segment corresponding to the target image frame matching the query information is retrieved from the database, and the cross-modal matching of text and image is realized. Then, the first generation module 830 generates the retrieval result for the query information according to the video data corresponding to the target video segment. In this way, through this cross-modal matching method based on the similarity calculation of the text feature vector and the image feature vector, it is effectively solved that due to the limited number of preset event tags, some video information in the database is difficult to be labeled, and then in the actual retrieval process, the query text not covered by the preset event tag cannot be matched to the corresponding video information, and finally misses detection and causes the problem of low accuracy of the retrieval result.
[0116] In some embodiments of the present application, the information retrieval device 800 in the embodiment of the present application may further include: The second acquisition module is used to obtain the event type corresponding to the query information, the event type includes a first event type or a second event type, the occurrence frequency of the first event type is higher than the occurrence frequency of the second event type, and the danger level of the first event type is lower than the danger level of the second event type, and the event type is determined by at least one of the following methods: based on a reference event tag matched with the query information, and based on keywords extracted from the query information.
[0117] The first determination module 820 may be specifically configured to determine whether there is a target video segment in the database according to a first similarity value between the first text feature vector of the event type and the query information and an image feature vector of each image frame in the database.
[0118] In some embodiments of the present application, the above-mentioned first determination module 820 can be specifically used to determine whether there is a first target image frame in the database according to the first similarity value of the first text feature vector of the query information and the image feature vector of each image frame in the database, when the event type includes a first event type, the first preset similarity value includes a first reference similarity value, the target image frame includes a first target image frame, and the first similarity value between the first text feature vector and the image feature vector of the first target image frame is greater than the first reference similarity value; and, in the case where there is a first target image frame in the database, determine whether there is a target video segment in the database.
[0119] In some embodiments of the present application, the first determining module 820 may be specifically used to: When the event type includes the second event type, the first preset similarity value includes the second reference similarity value, the target image frame includes the second target image frame, and the first similarity value between the first text feature vector and the image feature vector of the second target image frame is greater than the second reference similarity value, comparing the first text feature vector of the query information with the sound feature vector of each audio information in the database to obtain a second similarity value between the first text feature vector and the sound feature vector of each audio information; Determine whether there is a second target image frame in the database according to a first text feature vector of the query information and a first similarity value of an image feature vector of each image frame in the database; Determine whether there is target audio information in the database according to a second similarity value between the first text feature vector and the sound feature vector of each audio information; In the case that there is a second target image frame and / or target audio information in the database, the video segment corresponding to the second target image frame and / or target audio information is determined as the target video segment.
[0120] In some embodiments of the present application, the first determining module 820 may be specifically used to: In a case where the query information includes a location keyword, the event type includes a second event type, the first preset similarity value includes a second reference similarity value, the target image frame includes a second target image frame and the first text feature vector and a first similarity value of the image feature vector of the second target image frame is greater than the second reference similarity value, comparing the second text feature vector of the location keyword with the third text feature vector of each location tag in the database to obtain a third similarity value between the second text feature vector and the third text feature vector of each location tag; Determine whether there is a target location tag in the database according to a third similarity value between the second text feature vector and the third text feature vector of each location tag, the third similarity value between the second text feature vector and the third text feature vector of the target location tag is greater than the second reference similarity value, and the location tag is an identifier of location information of an event recorded in a video segment corresponding to the location tag; In the case where there is a second target image frame and / or a target position tag in the database, the video segment corresponding to the second target image frame and / or the target position tag is determined as the target video segment.
[0121] In some embodiments of the present application, the information retrieval device 800 may further include: A third acquisition module is used to acquire keywords of the query information and reference event tags corresponding to the keywords; A second determination module is used to search the video tag corresponding to the reference event tag in the database according to the reference event tag to determine whether there is a target video tag in the database, where the video tag includes a sample event tag and / or a sample event attribute tag; The third determination module is used to determine the video segment corresponding to the target video tag as the target video segment when there is a target video tag in the database.
[0122] Based on this, the first determination module 820 can be specifically used to determine whether there is a target video segment in the database based on the first similarity value between the first text feature vector of the query information and the image feature vector of each image frame in the database when there is no target video tag in the database.
[0123] In some embodiments of the present application, the first determining module 820 may be specifically used to: Obtaining a fourth similarity value between an event type corresponding to a reference event tag and a video tag corresponding to the reference event tag in a database, where the fourth similarity value is used to characterize a degree of similarity between the video tag and the video segment corresponding to the video tag, where the event type includes a first event type or a second event type, where the occurrence frequency of the first event type is higher than the occurrence frequency of the second event type, and where the danger level of the first event type is lower than the danger level of the second event type; According to the reference event tag, the event type and the fourth similarity value of the video tag corresponding to the reference event tag in the database, the video tag corresponding to the reference event tag in the database is retrieved to determine whether there is a target video tag in the database, and the fourth similarity value corresponding to the target video tag is greater than the second preset similarity value, and the second preset similarity value is determined based on the event type.
[0124] In some embodiments of the present application, the information retrieval device 800 may further include: A third receiving module is used to receive reference video information sent by an electronic device, the reference video information includes a video segment, a video identification tag, audio information, an image frame and a sample event attribute tag, the sample event attribute tag includes a time tag, a location tag, a video object tag and an alarm audio tag; A second generating module, used to generate a sample event label of the video segment according to the sample event attribute label; The third generation module is used to perform feature extraction and vector conversion on the image frame to generate an image feature vector of the image frame; The construction module is used to associate and store reference video information, sample event labels and image feature vectors of image frames to construct a database.
[0125] The various modules of the information retrieval device 800 provided in the embodiment of the present application can realize Figures 1 to 6 The functions of each step of the provided information retrieval method and its ability to achieve corresponding technical effects are described briefly and will not be repeated here.
[0126] The present application also provides an information retrieval device 900, which is applied to electronic devices such as Fig.10 As shown, the information retrieval device 900 includes: The first receiving module 910 is used to receive inquiry information input by the user; A first sending module 920 is used to send a search request to a cloud server, where the search request carries query information and is used to request the cloud server to search for search results for the query information in a database; The second receiving module 930 is used to receive a search result sent by the cloud server, where the search result is generated by the cloud server based on the video information of the target video segment, where the target video segment is determined by the cloud server based on a first similarity value between a first text feature vector of the query information and an image feature vector of each image frame in the database, where the target video segment includes a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value; The display module 940 is used to display the search results.
[0127] Therefore, by using cloud servers to perform database retrieval operations, electronic devices do not need to have strong local computing capabilities and large-capacity storage to process video data and perform complex text and image feature vector calculations. By handing over information retrieval tasks to cloud servers, they can fully utilize the advantageous resources of cloud servers to achieve faster and more accurate retrieval.
[0128] In some embodiments of the present application, the information retrieval device 900 may further include: A fourth acquisition module is used to acquire initial video information, where the initial video information includes an initial video segment, initial audio information, initial time information, and initial position information; A processing module is used to perform data preprocessing on the initial video information to obtain reference video information, where the reference video information includes a video segment, a video identification tag, audio information, an image frame, and a sample event attribute tag, where the sample event attribute tag includes a time tag, a location tag, a video object tag, and an alarm audio tag; The second sending module is used to send reference video information to the cloud server.
[0129] The various modules of the information retrieval device 900 provided in the embodiment of the present application can be implemented Figure 7 and Figure 8 The functions of each step of the provided information retrieval method and its ability to achieve corresponding technical effects are described briefly and will not be repeated here.
[0130] Fig.11 A schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application is shown.
[0131] The electronic device may include a processor 1101 and a memory 1102 storing computer program instructions.
[0132] Specifically, the processor 1101 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0133] The memory 1102 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 1102 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 1102 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 1102 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 1102 is a non-volatile solid-state memory.
[0134] In a particular embodiment, the memory 1102 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory 1102 includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the data processing method according to the first aspect of the present application.
[0135] The processor 1101 implements any one of the information retrieval methods in the above embodiments by reading and executing computer program instructions stored in the memory 1102 .
[0136] In one example, the electronic device may further include a communication interface 1103 and a bus 1110. Fig.11 As shown, the processor 1101, the memory 1102, and the communication interface 1103 are connected via a bus 1110 and communicate with each other.
[0137] The communication interface 1103 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0138] The bus 1110 includes hardware, software or both, coupling the components of the electronic device to each other. For example and not limitation, the bus may include an accelerated graphics port (AGP) or other graphics bus, an enhanced industry standard architecture (EISA) bus, a front-side bus (FSB), a hypertransport (HT) interconnect, an industry standard architecture (ISA) bus, an infinite bandwidth interconnect, a low pin count (LPC) bus, a memory bus, a microchannel architecture (MCA) bus, a peripheral component interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a serial advanced technology attachment (SATA) bus, a video electronics standard association local (VLB) bus or other suitable bus or a combination of two or more of these. Where appropriate, the bus 1110 may include one or more buses. Although the present application embodiment describes and illustrates a specific bus, the present application considers any suitable bus or interconnect.
[0139] The electronic device can execute the information retrieval method in the embodiment of the present application, thereby realizing the combination Figures 1 to 10 An information retrieval method and apparatus are described.
[0140] In addition, in combination with the information retrieval method in the above embodiment, the embodiment of the present application may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the information retrieval methods in the above embodiment is implemented. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as portable disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, etc.
[0141] In addition, in combination with the information retrieval method in the above embodiment, the embodiment of the present application may provide a computer program product for implementation. The program product is stored in a storage medium, and may specifically include a computer program or instruction, which implements any one of the information retrieval methods in the above embodiment when executed by a processor. The program product is executed by at least one processor to implement the various processes of the above data processing method embodiment, and can achieve the same technical effect, so it will not be repeated here to avoid repetition.
[0142] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0143] The functional blocks shown in the above structural block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier. "Machine-readable medium" may include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0144] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0145] Aspects of the present disclosure are described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0146] The above are only specific implementation methods of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. An information retrieval method, characterized in that: Applied to a cloud server, the method comprises: Obtain inquiry information; Determine whether there is a target video segment in the database according to a first similarity value between a first text feature vector of the query information and an image feature vector of each image frame in the database, the target video segment includes a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value; In the case where the target video segment exists in the database, a search result for the query information is generated according to the video information of the target video segment.
2. The information retrieval method according to claim 1, characterized in that: Before determining whether there is a target video segment in the database based on the first text feature vector of the query information and the first similarity value of the image feature vector of each image frame in the database, the method further includes: Acquire an event type corresponding to the query information, the event type including a first event type or a second event type, the occurrence frequency of the first event type is higher than the occurrence frequency of the second event type, and the danger level of the first event type is lower than the danger level of the second event type, and the event type is determined in at least one of the following ways: based on a reference event tag matched by the query information, or based on a keyword extracted from the query information; The determining whether there is a target video segment in the database according to the first text feature vector of the query information and the first similarity value of the image feature vector of each image frame in the database comprises: It is determined whether there is a target video segment in the database according to the first similarity value between the event type and the first text feature vector of the query information and the image feature vector of each image frame in the database.
3. The information retrieval method according to claim 2, characterized in that: The event type includes a first event type, the first preset similarity value includes a first reference similarity value, the target image frame includes a first target image frame, and a first similarity value between the first text feature vector and the image feature vector of the first target image frame is greater than the first reference similarity value; The determining whether there is a target video segment in the database according to the first text feature vector of the event type and the query information and the first similarity value of the image feature vector of each image frame in the database comprises: Determining whether the first target image frame exists in the database according to a first text feature vector of the query information and a first similarity value of an image feature vector of each image frame in the database; In a case where the first target image frame exists in the database, it is determined that there is a target video segment in the database.
4. The information retrieval method according to claim 2, characterized in that: The event type includes a second event type, the first preset similarity value includes a second reference similarity value, the target image frame includes a second target image frame, and a first similarity value between the first text feature vector and the image feature vector of the second target image frame is greater than the second reference similarity value; The determining whether there is a target video segment in the database according to the first text feature vector of the event type and the query information and the first similarity value of the image feature vector of each image frame in the database comprises: Comparing the first text feature vector of the inquiry information with the sound feature vector of each audio information in the database to obtain a second similarity value between the first text feature vector and the sound feature vector of each audio information; Determining whether the second target image frame exists in the database according to a first text feature vector of the query information and a first similarity value of an image feature vector of each image frame in the database; determining whether there is target audio information in the database according to a second similarity value between the first text feature vector and the sound feature vector of each audio information, wherein the second similarity value between the first text feature vector and the sound feature vector of the target audio information is greater than the second reference similarity value; In a case where the second target image frame and / or the target audio information exists in the database, a video segment corresponding to the second target image frame and / or the target audio information is determined as the target video segment.
5. The information retrieval method according to claim 2, characterized in that: The query information includes a location keyword, the event type includes a second event type, the first preset similarity value includes a second reference similarity value, the target image frame includes a second target image frame, and a first similarity value between the first text feature vector and the image feature vector of the second target image frame is greater than the second reference similarity value; The determining whether there is a target video segment in the database according to the first text feature vector of the event type and the query information and the first similarity value of the image feature vector of each image frame in the database comprises: Comparing the second text feature vector of the location keyword with the third text feature vector of each location tag in the database to obtain a third similarity value between the second text feature vector and the third text feature vector of each location tag; Determine whether there is a target location tag in the database according to a third similarity value between the second text feature vector and the third text feature vector of each location tag, the third similarity value between the second text feature vector and the third text feature vector of the target location tag is greater than the second reference similarity value, and the location tag is an identifier of location information of an event recorded in a video segment corresponding to the location tag; In a case where the second target image frame and / or the target position tag exist in the database, a video segment corresponding to the second target image frame and / or the target position tag is determined as the target video segment.
6. The information retrieval method according to any one of claims 1 to 4, characterized in that: Before determining whether there is a target video segment in the database based on the first text feature vector of the query information and the first similarity value of the image feature vector of each image frame in the database, the method further includes: Acquire keywords of the query information and reference event tags corresponding to the keywords; According to the reference event tag, searching the database for a video tag corresponding to the reference event tag to determine whether there is a target video tag in the database, wherein the video tag includes a sample event tag and / or a sample event attribute tag; In a case where there is a target video tag in the database, determining a video segment corresponding to the target video tag as the target video segment; The determining whether there is a target video segment in the database according to the first text feature vector of the query information and the first similarity value of the image feature vector of each image frame in the database comprises: In the case that the target video tag does not exist in the database, it is determined whether the target video segment exists in the database according to the first text feature vector of the query information and the first similarity value of the image feature vector of each image frame in the database.
7. The information retrieval method according to claim 6, characterized in that: The step of searching the database for a video tag corresponding to the reference event tag according to the reference event tag to determine whether the database contains a target video tag comprises: Obtaining a fourth similarity value between an event type corresponding to the reference event tag and a video tag corresponding to the reference event tag in the database, wherein the fourth similarity value is used to characterize a similarity between the video tag and the video segment corresponding to the video tag, wherein the event type includes a first event type or a second event type, wherein an occurrence frequency of the first event type is higher than an occurrence frequency of the second event type, and a danger level of the first event type is lower than a danger level of the second event type; According to the reference event tag, the event type and the fourth similarity value of the video tag corresponding to the reference event tag in the database, the video tag corresponding to the reference event tag in the database is retrieved to determine whether there is a target video tag in the database, and the fourth similarity value corresponding to the target video tag is greater than a second preset similarity value, and the second preset similarity value is determined based on the event type.
8. The information retrieval method according to any one of claims 1 to 4, characterized in that: The method further comprises: Receiving reference video information sent by an electronic device, the reference video information including a video segment, a video identification tag, audio information, an image frame, and a sample event attribute tag, the sample event attribute tag including a time tag, a location tag, a video object tag, and an alarm audio tag; Generating a sample event label of the video segment according to the sample event attribute label; Performing feature extraction and vector conversion on the image frame to generate an image feature vector of the image frame; The database is constructed by associating and storing the reference video information, the sample event label and the image feature vector of the image frame.
9. An information retrieval method, characterized in that: Applied to electronic equipment, the method comprises: Receive inquiry information input by the user; Sending a search request to a cloud server, the search request carrying the query information, the search request being used to request the cloud server to search a database for a search result for the query information; receiving a retrieval result sent by the cloud server, wherein the retrieval result is generated by the cloud server based on video information of a target video segment, wherein the target video segment is determined by the cloud server according to a first similarity value between a first text feature vector of the query information and an image feature vector of each image frame in a database, wherein the target video segment includes a target image frame, and a first similarity value between the first text feature vector and the image feature vector of the target image frame is greater than a first preset similarity value; The search results are displayed.
10. The method according to claim 9, characterized in that The method further comprises: Acquire initial video information, where the initial video information includes an initial video segment, initial audio information, initial time information, and initial position information; Performing data preprocessing on the initial video information to obtain reference video information, wherein the reference video information includes a video segment, a video identification tag, audio information, an image frame, and a sample event attribute tag, wherein the sample event attribute tag includes a time tag, a location tag, a video object tag, and an alarm audio tag; The reference video information is sent to a cloud server.
Citation Information
Patent Citations
Video tracing method and device, computer equipment and storage medium
CN114554302A
Model reasoning service device, model pre-training device and video clip retrieval method
CN119719417A
Content processing method and apparatus, computer device, and storage medium
WO2021223567A1