A method for constructing spatiotemporal knowledge graph
By building a space-time knowledge graph based on device specification description and user behavior logs, the problem of low information islands and intelligent service levels in smart home systems is solved, and a higher level of intelligent and personalized service is achieved.
Patent Information
- Application Number
- CN202411975364.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing smart home systems have problems with information silos between devices and low level of intelligent services, resulting in high complexity when managing and controlling devices and insufficient intelligent services.
The space-time knowledge graph construction method based on device specification description and user behavior log is adopted. The device specification description knowledge graph is constructed through image-text matching and text structure extraction, and the space-time knowledge graph is constructed based on user behavior logs to optimize the equipment collaborative work and intelligent services.
It significantly improves the intelligence level of the system, optimizes the collaborative work between devices, provides more personalized and intelligent home services, and improves the user experience.
Smart Images

Figure CN119377420B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart home technology, and specifically, relates to a method for constructing a spatiotemporal knowledge graph based on device specifications and user behavior logs. Background Art
[0002] In the field of smart home, with the development of Internet of Things technology, smart home devices have gradually become popular, and family life has become more and more convenient. Smart home systems usually include smart lighting, smart temperature control, security monitoring, voice assistants, etc. These devices are connected to each other through the Internet to provide personalized services and improve the quality of life.
[0003] However, the current smart home system still has the following two problems in practical application. First, the devices and services of existing smart home systems are often independent of each other and lack an effective collaborative working mechanism. This leads to information islands between devices, making it more complex and challenging for users to manage and control these devices. Furthermore, although some systems can be operated through voice assistants, these assistants often lack a deep understanding of the home environment and device status and cannot provide truly intelligent services. Users need to manually input or adjust settings, which to some extent reduces the intelligence level of the system.
[0004] However, with the continuous advancement of knowledge graph technology and large language model technology, these technologies can effectively solve the problems of information islands and low intelligence level mentioned above to varying degrees. As a graphical knowledge representation method, knowledge graph can integrate multi-dimensional data from different devices and services to form a comprehensive and structured knowledge network, thereby improving the system's understanding of the home environment and service quality. Despite this, there are still deficiencies in natural language processing, context understanding and multi-task processing. However, these problems can be solved by large language models. Large language models (such as GPT-4) have powerful natural language understanding and generation capabilities, can extract valuable information from large amounts of text data, and perform intelligent reasoning. Therefore, combining the information integration and knowledge reasoning capabilities of knowledge graphs with the natural language processing and context understanding capabilities of large language models can effectively solve the problems of information islands, insufficient device collaboration, and low intelligent service levels in existing smart home systems, bringing new breakthroughs to the development of smart homes. However, constructing knowledge graphs in the smart home field faces difficulties and challenges such as multiple data sources, low data quality, and small data volumes. Summary of the invention
[0005] The purpose of the present invention is to provide a method for constructing a spatiotemporal knowledge graph based on device specifications and user behavior logs, which can infer user habits and predict user behaviors through the basic knowledge existing in the scene and the dynamically generated user behavior logs, thereby providing more intelligent, proactive, smart and personalized services, and significantly improving user experience.
[0006] The present invention is implemented by the following technical solutions:
[0007] A method for constructing a spatiotemporal knowledge graph is proposed, including:
[0008] Parsing the device specification document to obtain its textual representation and visual representation;
[0009] Perform image-text matching and build a specification ontology model;
[0010] Generate prompts using the specification ontology model, concatenate them with the text representation and feed them into the large language model for text structure extraction to obtain the device specification knowledge graph;
[0011] Extract candidate log pairs based on user behavior log information;
[0012] Determine the joint functional similarity between devices in the candidate log pairs based on the constructed specification knowledge graph;
[0013] Construct a spatiotemporal knowledge graph based on the joint similarity of features between candidate log pairs.
[0014] In some embodiments of the present invention, after constructing the spatiotemporal knowledge graph, the method further includes:
[0015] The scaling factor is used to adjust the numerical scale of the spatiotemporal knowledge graph as a whole and update it to obtain the final spatiotemporal knowledge graph.
[0016] In some embodiments of the present invention, parsing the specification document to obtain its text representation and visual representation includes:
[0017] For text information, the word embedding matrix from the pre-trained model is used to initialize the text embedding and incorporate the position information contained in the text information;
[0018] For visual information, the output features of the CNN visual encoder are mapped to incorporate the position information contained in the visual information;
[0019] For layout information, use OCR to identify the information contained in the bounding box and merge the horizontal coordinate information with the vertical coordinate information;
[0020] The processed text embedding, visual embedding and layout embedding are fused and sent to the multimodal model to obtain the textual representation and visual representation of the specification document.
[0021] In some embodiments of the present invention, performing image-text matching specifically includes:
[0022] Align the textual and visual representations of the specification document, calculate the region-text similarity, and obtain the image-text similarity through maximum pooling;
[0023] The relative position of image and text is embedded into the image and text similarity to obtain the image and text fusion similarity, and the image and text similarity is obtained after LSE pooling;
[0024] Image-text matching with relative position information in specification documents is performed based on image-text similarity ranking.
[0025] In some embodiments of the present invention, extracting candidate log pairs based on user behavior log information specifically includes:
[0026] Constructing a device standard time threshold matrix based on the time intervals set between associated devices;
[0027] Based on the device standard time threshold matrix, candidate log pairs that meet the threshold are extracted from the user behavior log information.
[0028] In some embodiments of the present invention, in determining the joint function similarity between devices in a candidate log pair, a calculation function of the joint function similarity is expressed as:
[0029] ;
[0030] in, Represents the i-node in the knowledge graph; Indicates the functional text description corresponding to the node; For text embedding representation using RoBERTa, is the hidden layer representation, is the hidden state of the i-th layer, The last layer of hidden states.
[0031] In some embodiments of the present invention, constructing a spatiotemporal knowledge graph based on the functional similarity between candidate log pairs includes:
[0032] Set the function joint similarity threshold;
[0033] Increase the relationship frequency between devices corresponding to the log pairs whose joint functional similarity is greater than the threshold by one;
[0034] The relationship between the devices corresponding to the daily systems whose joint function similarity is less than the threshold is set to the maximum value.
[0035] In some embodiments of the present invention, the numerical scale of the spatiotemporal knowledge graph is adjusted overall using a scaling factor, including:
[0036] Traverse the relationships in the spatiotemporal knowledge graph and keep the maximum value of the relationship in the max variable;
[0037] For the relationship values that are greater than the set threshold, the relationship values are set to 0 and the relationship is saved in the Invalid array;
[0038] After the traversal is completed, if the maximum value of the relationship is greater than the set value, the scaling factor is used to adjust the size of the overall value;
[0039] Output the adjusted spatiotemporal knowledge graph.
[0040] Compared with the prior art, the advantages and positive effects of the present invention are as follows: the present invention proposes a method for constructing a spatiotemporal knowledge graph, and for each device in the smart home space, the specification document is respectively subjected to image text matching and text structured extraction, and the specification knowledge graph of each device is constructed, and then the information contained in the device specification knowledge graph is used to filter out valid log information pairs in the smart home space, and the smart home spatiotemporal knowledge graph is constructed according to the joint functional similarity between the devices contained in the log pairs. The present invention utilizes the rich and effective device information contained in the specification documents of the devices in the smart home space, as well as the timing information and user behavior patterns contained in the user behavior logs, and combines the knowledge graph and the large language model to understand and process the specification documents and user behavior log files, thereby constructing a comprehensive, dynamic and time-space evolving smart home domain knowledge graph, which can not only optimize the collaborative work between devices, but also significantly improve the intelligence level of the system, enabling users to more conveniently manage and control the home environment, and realize more personalized and intelligent home services.
[0041] Other features and advantages of the present invention will become more apparent after reading the detailed description of the embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0043] Figure 1 This is a schematic diagram of the steps of the spatiotemporal knowledge graph construction method proposed in the present invention;
[0044] Figure 2 This is a schematic diagram of the spatiotemporal knowledge graph construction process proposed in the present invention;
[0045] Figure 3 The following is a schematic diagram of the process of parsing the specification document in the present invention;
[0046] Figure 4 This is a schematic diagram of the process of image-text matching in the present invention;
[0047] Figure 5 A schematic diagram of a specification ontology model constructed in the present invention;
[0048] Figure 6 It is a schematic diagram of the text structure extraction process in the present invention;
[0049] Figure 7 An embodiment of the present invention for generating prompts according to a specification ontology;
[0050] Figure 8 This is a schematic diagram of the process of extracting candidate log pairs in the present invention;
[0051] Fig. 9 This is a schematic diagram of the process of constructing a spatiotemporal knowledge graph in the present invention;
[0052] Fig.10 An example of a specification document of a device in an embodiment of the present invention;
[0053] Fig.11 In the embodiment of the present invention, according to Fig.10 The specification knowledge graph generated by the specification document shown is shown;
[0054] Fig.12 This is a schematic diagram of a user behavior log file in an embodiment of the present invention;
[0055] Fig.13 Schematic diagram of the original spatiotemporal knowledge graph constructed in an embodiment of the present invention;
[0056] Fig.14 Schematic diagram of a standard time threshold matrix set in an embodiment of the present invention;
[0057] Fig.15 This is an illustration of judging functional similarity of two devices in a log pair according to two constructed specification knowledge graphs in an embodiment of the present invention;
[0058] Fig.16 This is an example of a spatiotemporal knowledge graph after overall adjustment by a scaling factor in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] In the description of the present invention, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0061] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0062] The terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0063] The specification documents and user behavior log files in the smart home space contain a lot of effective information. Among them, the specification contains rich and effective device information in the home space, and the user behavior log contains rich time series information and user behavior patterns. This information plays a key role in building a comprehensive, dynamic and efficient reasoning knowledge graph. At the same time, combining the knowledge graph with the large language model can better understand and process the specification documents and user behavior log files in the home space, thereby building a comprehensive, dynamic and time-evolving smart home domain knowledge graph. This can not only optimize the collaborative work between devices, but also significantly improve the intelligence level of the system, allowing users to manage and control the home environment more conveniently and achieve more personalized and intelligent home services.
[0064] For example, at present, most of the scene generation in the home relies on manual input from users, and some rely on mobile phone apps to trigger the generation of scenes. In the event that the user comes home from get off work to cook dinner, the current scene generation still relies on the user to manually define the time to go home and turn on the lights and air conditioners, or to bind the smart door lock opening event with turning on the living room lights and turning on the air conditioner, and manually set the brightness of the lights and the mode temperature of the air conditioner. The kitchen lights can be turned on by voice control or manually controlled by the user, and the cooking equipment can be manually started after entering the kitchen, or started in advance through the mobile phone app.
[0065] After using the constructed spatiotemporal knowledge graph of smart home space, the scenario of users coming home from get off work to cook dinner can be described as follows: the system detects the user's approach to the community through the user's mobile phone positioning, and judges whether to close the windows in advance, turn on the air conditioner, and the mode and temperature of the air conditioner through the real-time temperature and humidity data from the temperature and humidity sensor in the knowledge graph, and judges whether to turn on the cooking equipment. When the user opens the smart door lock, the real-time brightness from the illuminance sensor in the knowledge graph determines whether to automatically turn on the corresponding lamps and automatically adjust the light brightness. The process of the user passing through the corridor to the kitchen determines the user's direction through the changes in the information of the people in the corridor and the kitchen, and then judges whether to turn on the lights and the light brightness in combination with the value of the illuminance sensor in the kitchen.
[0066] In general, the present invention aims to construct such a spatiotemporal knowledge graph, which can infer users' habits and predict users' behaviors through the basic knowledge existing in the scene and dynamically generated user behavior logs, and then provide more intelligent, proactive, smart and personalized services, significantly improving the user experience.
[0067] The present invention proposes a method for constructing a spatiotemporal knowledge graph based on device specifications and user behavior logs, such as Figure 1 and Figure 2 As shown, including:
[0068] In the first stage, a knowledge graph of specifications of each device in the smart home space is constructed.
[0069] In this stage, for all devices in the smart home space, a specification knowledge graph is built for each device. That is, one knowledge graph corresponds to a device object, which contains all the information in the current specification document object, mainly including the entity object corresponding to the specification document, the directories at all levels in the specification document, the detailed content under the directory, and the pictures corresponding to some of the content.
[0070] It is mainly divided into the following four parts:
[0071] S1: Parse the specification document to obtain its textual representation and visual representation.
[0072] For text information, the word embedding matrix from the pre-trained model RoBERTa (Robustly Optimized BERTApproach) is used to initialize the text embedding, and then the position information contained in the text information is incorporated. The text embedding is represented as:
[0073] ;
[0074] in Indicates text snippets, express Semantic word embeddings, Indicates The one-dimensional position embedding represented by the token index of the text fragment, Indicates The 2D position embedding represented by the bounding box coordinates of each text snippet.
[0075] For visual information, the output features of the CNN (Convolutional Neural Networks) visual encoder are mapped and then integrated into the position information contained in the visual information. The visual embedding is expressed as:
[0076] ;
[0077] in represents the input image, Indicates The feature maps of the output images are average pooled and flattened after the image is input into the visual backbone network. represents a linear projection layer that unifies the dimensions of the visual embedding sequence with the text embedding, Indicates The one-dimensional position embedding represented by the label index of the visual segment. In the present invention, all visual information is attached to the visual segment [C]. is the semantic visual embedding of the current image.
[0078] For layout information, use OCR (Optical Character Recognition) to identify the information contained in the bounding box and merge the horizontal coordinate information with the vertical coordinate information. The layout embedding is represented as:
[0079] ;
[0080] ;
[0081] in, Indicates the coordinate position information of the current text or visual mark box. The embedding layer representing the x-axis features, The embedding layer representing the y-axis features, It means that the two embedding layers will be connected along the dimension.
[0082] As mentioned above, the preprocessed text embedding, visual embedding, and layout embedding are fused and sent to the multimodal Transformer to obtain the text representation and visual representation of the specification document, as shown in Figure 3 shown.
[0083] S2: Perform image-text matching.
[0084] The textual and visual representations of the specification document are aligned, and then the region-text similarity is calculated, and the image-text similarity is obtained after maximum pooling.
[0085] The image-text relative position is embedded into the image-text similarity to obtain the image-text fusion similarity.
[0086] Assume that the object in the specification document Expressed as ,in Indicates the page number of the current content. , Indicates the coordinates of the upper right corner of the current object. , Represents the width and height of the current object. Then, the relative position embedding is expressed as:
[0087] ;
[0088] in, is the view matrix, which is used to transform the relative position information from the local space to the semantic space.
[0089] Finally, the image-text similarity is obtained through LSE pooling, and the image-text matching of the fused relative position information in the specification document is performed based on this similarity ranking, such as Figure 4 shown.
[0090] S3: Construct specification ontology model.
[0091] The ontology of the specification document is defined to help the large language model enhance its understanding of the content and structure of the device specification document and more effectively parse and extract structured text information from the document. The ontology of the specification document includes but is not limited to: Figure 5 The content shown (only the text that can help the large language model to extract text information in a structured manner is shown, and irrelevant ontology attributes are not shown).
[0092] S4: Generate prompts using the specification ontology model and feed them into the large language model for text structured extraction to obtain the device specification knowledge graph.
[0093] The above specification ontology model is used to construct prompts, which are combined with the text representation in the specification manual and sent to the large language model to prompt the large language model to extract structured information, and finally obtain structured text information, such as Figure 6 shown.
[0094] An example of an input generated is Figure 7 shown.
[0095] In the second stage, the spatiotemporal knowledge graph of smart home space is constructed based on the specification knowledge graph.
[0096] In this step, the user behavior log information is used to build the smart home space-time knowledge graph. The transformation of each device state recorded in the user behavior log and the multiple device specification knowledge graphs generated in the previous stage are used to infer whether new relationships have been generated between devices. Then, the relationship values are controlled as a whole to eliminate irrelevant event connections caused by noise behaviors, and finally the smart home space-time knowledge graph is built. The construction process is as follows: Figure 8 As shown, it is mainly divided into the following four parts:
[0097] S5: Extract candidate log pairs based on user behavior logs.
[0098] First, we build a device standard time threshold matrix based on the reasonable time intervals between associated devices. , and extract candidate log pairs that meet the time threshold in the user behavior log information L according to the matrix. The user behavior log information L includes the device name, device number, status, and time of state change.
[0099] S6: Determine the joint functional similarity between devices in the candidate log pairs based on the constructed specification knowledge graph.
[0100] With the help of the constructed specification knowledge graph, the functional association between the devices in the candidate log pairs is determined. The similarity calculation function of the functional association is expressed as:
[0101] ;
[0102] in, Represents the i-node in the knowledge graph; Indicates the functional text description corresponding to the node; For text embedding representation using RoBERTa, is the hidden layer representation, is the hidden state of the i-th layer, The last hidden state. "." represents any content.
[0103] S7: Construct a spatiotemporal knowledge graph based on the functional similarity between candidate log pairs.
[0104] Based on the joint function similarity between the candidate log pairs, we determine whether the current log pairs have a correlation. We increase the frequency of the relationship between the devices corresponding to the log pairs with joint functions (joint function similarity is greater than the threshold) by one, and set the relationship between the devices corresponding to the log pairs with antagonistic functions (joint function similarity is less than the threshold) to the maximum value, thus completing the construction of the preliminary spatiotemporal knowledge graph.
[0105] S8: Use the scaling factor to adjust the numerical scale of the spatiotemporal knowledge graph as a whole and update it to obtain the final spatiotemporal knowledge graph.
[0106] For the created spatiotemporal knowledge graph, it represents the relationship (concurrency level) between devices in the home scenario. Each node corresponds to a device object. To address the problem of value crossing the boundary and noise logs generated by accidental user behavior leading to connections between irrelevant devices, the present invention uses a scaling factor to adjust the value scale as a whole to eliminate value crossing the boundary and connections between irrelevant devices, such as Fig. 9 As shown, including:
[0107] 1. Traverse the relationships in the spatiotemporal knowledge graph and keep the maximum value of the relationship in the max variable.
[0108] 2. For relationship values greater than the set threshold, set the relationship value to 0 and save the relationship in the Invalid array.
[0109] 3. After the traversal is completed, if the maximum value of the relationship is greater than the set value (in the example, 100, which is the limit set according to actual requirements), use the scaling factor to adjust the size of the overall value.
[0110] 4. Output the adjusted spatiotemporal knowledge graph.
[0111] The spatiotemporal knowledge graph construction method proposed in the present invention is described in detail below using a specific embodiment.
[0112] like Fig.10The text embedding, visual embedding, and layout embedding of an example of a device specification document shown in the figure are input into a multimodal converter to obtain text representation and visual representation, the similarity of the two is then calculated, and the relative position embedding of the text segment and the image is integrated after maximum pooling and then LSE pooling is performed to obtain image-text similarity. Image-text matching is performed based on the ranking of the similarity, and the text representation is combined with the prompts generated by the constructed specification ontology model and sent to the large language model for text structured extraction. The structured text and one-to-one image-text pairs constitute the complete specification document knowledge graph of the current specification document.
[0113] like Fig.11 In the visualized knowledge graph, red nodes represent the device entities corresponding to the specification document, dark blue nodes represent the first-level directory of the specification document, yellow nodes represent the second-level directory, light blue nodes represent text segments, and green nodes represent images in the specification document. There are content connections between these nodes, forming a specification knowledge graph.
[0114] Next, the spatiotemporal knowledge graph is updated according to the spatiotemporal knowledge graph, user behavior log, standard time threshold matrix, and similarity threshold. If the spatiotemporal knowledge graph does not exist, the spatiotemporal knowledge graph creation operation is executed. If there is Fig.12 As shown in the following example, Fig.13 The spatiotemporal knowledge graph KG shown in Fig.14 The standard time threshold matrix S shown in the figure can screen out candidate log pairs that meet the standard time threshold. The two devices in the log pair are judged for their functional union according to the two specification knowledge graphs constructed in the previous stage. If the functions are united, the relationship frequency between the corresponding device nodes is increased by 1. If the functions are antagonistic, the corresponding relationship is set to the maximum value. After the above steps, a preliminary spatiotemporal knowledge graph can be obtained, as shown in Fig.15 As shown in the figure, a scaling factor is then used to adjust the value threshold in the spatiotemporal knowledge graph as a whole to eliminate the value crossing the boundary and the noise time, and finally the following is obtained: Fig.16 The spatiotemporal knowledge graph shown.
[0115] It should be pointed out that the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for constructing a spatiotemporal knowledge graph, characterized in that: include: Parsing the device specification document to obtain its text representation and visual representation; Perform image-text matching and construct a specification ontology model; the specification ontology model is to define the ontology of the specification document, and the definition is used to help the large language model enhance its understanding of the content and structure of the device description document, so as to help the large language model to extract text information in a structured manner; Generate prompts using the specification ontology model, concatenate them with the text representation and feed them into the large language model for text structure extraction to obtain the device specification knowledge graph; Extract candidate log pairs based on user behavior log information; Based on the constructed specification knowledge graph, the joint functional similarity between devices in the candidate log pairs is determined; the calculation function of the joint functional similarity is expressed as: ;in, Represents the i-node in the knowledge graph; Indicates the functional text description corresponding to the node; For text embedding representation using RoBERTa, is the hidden layer representation, is the hidden state of the i-th layer, The last layer of hidden state; Construct a spatiotemporal knowledge graph based on the joint similarity of features between candidate log pairs.
2. The method for constructing a spatiotemporal knowledge graph according to claim 1, characterized in that: After constructing the spatiotemporal knowledge graph, the method further includes: The scaling factor is used to adjust the numerical scale of the spatiotemporal knowledge graph as a whole and update it to obtain the final spatiotemporal knowledge graph.
3. The method for constructing a spatiotemporal knowledge graph according to claim 1, characterized in that: Parse the specification document to obtain its textual and visual representations, including: For text information, the word embedding matrix from the pre-trained model is used to initialize the text embedding and incorporate the position information contained in the text information; For visual information, the output features of the CNN visual encoder are mapped to incorporate the position information contained in the visual information; For layout information, use OCR to identify the information contained in the bounding box and merge the horizontal coordinate information with the vertical coordinate information; The processed text embedding, visual embedding and layout embedding are fused and sent to the multimodal model to obtain the textual representation and visual representation of the specification document.
4. The method for constructing a spatiotemporal knowledge graph according to claim 1, characterized in that: Image-text matching specifically includes: Align the textual and visual representations of the specification document, calculate the region-text similarity, and obtain the image-text similarity through maximum pooling; The relative position of image and text is embedded into the image and text similarity to obtain the image and text fusion similarity, and the image and text similarity is obtained after LSE pooling; Image-text matching with relative position information in specification documents is performed based on image-text similarity ranking.
5. The method for constructing a spatiotemporal knowledge graph according to claim 1, characterized in that: Extract candidate log pairs based on user behavior log information, including: Constructing a device standard time threshold matrix based on the time intervals set between associated devices; Based on the device standard time threshold matrix, candidate log pairs that meet the threshold are extracted from the user behavior log information.
6. The method for constructing a spatiotemporal knowledge graph according to claim 1, characterized in that: Construct a spatiotemporal knowledge graph based on the functional similarity between candidate log pairs, including: Set the function joint similarity threshold; The relationship frequency between the devices corresponding to the log pairs whose joint functional similarity is greater than the threshold is increased by one; The relationship between devices corresponding to the log pairs whose joint functional similarity is less than the threshold is set to the maximum value.
7. The method for constructing a spatiotemporal knowledge graph according to claim 2, characterized in that: Use scaling factors to adjust the numerical scale of the spatiotemporal knowledge graph as a whole, including: Traverse the relationships in the spatiotemporal knowledge graph and keep the maximum value of the relationship in the max variable; For the relationship values that are greater than the set threshold, the relationship values are set to 0 and the relationship is saved in the Invalid array; After the traversal is completed, if the maximum value of the relationship is greater than the set value, the scaling factor is used to adjust the size of the overall value; Output the adjusted spatiotemporal knowledge graph.
Citation Information
Patent Citations
Visual question-answering system based on multi-modal knowledge graph and large language model
CN118394906A
Terminal equipment-oriented intelligent scene generation method and smart home system
CN118444620A