An evidence screening method and system based on a large language model
By using a large language model to perform multimodal fusion of text, image, and audio evidence, an evidence chain timeline is constructed, solving the problem of time-consuming and labor-intensive manual sorting and achieving efficient and accurate evidence timeline organization.
Patent Information
- Application Number
- CN202510226095.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing technologies rely on manual sorting of fragmented narratives in case handling, resulting in high human resource consumption and a lack of creativity, and failing to efficiently organize the timeline of events.
By using a large language model to perform multimodal fusion of text, image and audio evidence, and by recognizing time nodes and constructing unit timelines, an evidence chain timeline and a time sequence evidence graph are generated, enabling the temporal order analysis and arrangement of evidence.
It achieves efficient integration of multiple evidence modalities, reduces human resource consumption, improves the accuracy and efficiency of evidence sorting, and generates a chain of evidence in chronological order.
Smart Images

Figure CN120148042B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of case processing, and in particular to an evidence screening method and system based on a large language model. BACKGROUND
[0002] With the continuous improvement of residents' legal consciousness and the increasing number of various disputes, the demand for case processing is also growing, and correspondingly, in the scenes of judicial inquiry, case investigation or psychological assessment, the timeline of the event is often fuzzy and the logic is broken.
[0003] At present, the traditional method relies on manual time sorting of scattered narratives.
[0004] The above method can realize the time sorting of scattered narratives, but it consumes human resources and energy, and when the narrative content is long, a large amount of human resources is needed, and only text recognition is needed, which is not a creative labor, and can be completed by a large language model, so how to apply a large language model to the analysis of the timeline to reduce unnecessary human resource consumption has become a problem to be solved. SUMMARY
[0005] The present application provides an evidence screening method based on a large language model and a computer readable storage medium, which mainly aims to realize the analysis and arrangement of evidence of different modalities in chronological order to assist in combing evidence.
[0006] To achieve the above purpose, the present application provides an evidence screening method based on a large language model, which comprises:
[0007] receiving an initial text evidence set, an initial image evidence set and an initial audio evidence set;
[0008] obtaining a to-be-processed text set based on the initial text evidence set, the initial image evidence set and the initial audio evidence set, wherein the to-be-processed text set comprises a text evidence set, an image text set and an audio text set;
[0009] extracting the to-be-processed text set from the to-be-processed text set in turn, and correcting the to-be-processed text set by using a pre-constructed natural language model to obtain an optimized text set;
[0010] extracting the optimized text from the optimized text set in turn, and performing the following operations on the optimized text:
[0011] identifying a time node set corresponding to the optimized text, and constructing a unit time axis corresponding to the optimized text based on the time node set, wherein the time node set comprises zero, one or more time nodes;
[0012] By summarizing the unit timelines, we obtain the unit timeline set corresponding to the optimized text set. Based on the unit timeline set, we construct the evidence chain timeline, which includes multiple evidence time nodes.
[0013] A temporal evidence graph is constructed based on the evidence chain timeline and an optimized text set to complete the evidence screening based on a large language model.
[0014] Optionally, obtaining the text set to be processed based on the initial text evidence set, the initial image evidence set, and the initial audio evidence set includes:
[0015] Initial image evidence is extracted sequentially from the initial image evidence set to obtain the target processing image. Based on the target processing image and the pre-constructed graph identifier structure, an image identifier for the target processing image is generated. The image identifier is used to perform an identifier operation on the target processing image. The identified target processing images are summarized to obtain an identifier image set. An image text set is obtained based on the identifier image set. The image text set includes multiple image texts, and each image text corresponds one-to-one with the image identifier of the identifier image.
[0016] The initial audio evidence set is transformed into a speech text set using a pre-constructed speech identifier structure. The speech text set includes multiple speech texts, and each speech text corresponds one-to-one with a speech identifier.
[0017] The voice text set, image text set, and initial text evidence set are combined to obtain the text set to be processed.
[0018] Optionally, the step of generating image tags for the target processed image based on the target processed image and a pre-constructed graph tagging structure, performing tagging operations on the target processed image using the image tags, and summarizing the tagged target processed images to obtain a tagged image set, includes:
[0019] If there is no preset text region in the target image, the target image is identified as a textless image, and a textless image identifier is generated based on the textless image and the image identifier structure. The textless image identifier is then used to perform an identifier operation on the textless image to obtain the identifier image.
[0020] Otherwise, the target processed image is cropped based on the text region to obtain a target text image set, wherein the target text image set includes one or more target text images;
[0021] Perform the following operations on a target text image or a target text image:
[0022] performing a convolution operation on the target text image by using a pre-constructed Gaussian filter to obtain a first smoothed image, performing an edge enhancement operation on the first smoothed image to obtain a text edge enhanced image, performing a background removal operation on the text edge enhanced image to obtain a text enhanced image, performing a text recognition operation based on the text enhanced image to obtain a target image text;
[0023] obtaining a target image text set based on the target image text set and the icon recognition structure, and performing an identification operation on the target processing image by using the text icon recognition to obtain an identified image;
[0024] obtaining an identified image set by aggregating the identified images.
[0025] Optionally, the performing a convolution operation on the target text image by using a pre-constructed Gaussian filter to obtain a first smoothed image, performing an edge enhancement operation on the first smoothed image to obtain a text edge enhanced image, performing a background removal operation on the text edge enhanced image to obtain a text enhanced image, includes:
[0026] generating a Gaussian filter by using a pre-constructed two-dimensional Gaussian kernel, and performing a convolution operation on the target text image by using the Gaussian filter to obtain a first smoothed image, wherein the convolution formula is as follows:
[0027]
[0028] P=G*M
[0029] wherein x represents a horizontal coordinate of an image coordinate system corresponding to the target text image, y represents a vertical coordinate of the image coordinate system corresponding to the target text image, G(x, y) represents a two-dimensional Gaussian kernel, G represents a Gaussian filter constructed based on a Gaussian filter function, P represents the first smoothed image, and M represents the target text image.
[0030] obtaining a text edge enhanced image based on the first smoothed image and a pre-constructed edge enhancement method;
[0031] obtaining an optimized image based on the text edge enhanced image, and generating a text enhanced image according to the optimized image and a pre-constructed decoder.
[0032] Optionally, the obtaining a text edge enhanced image based on the first smoothed image and a pre-constructed edge enhancement method includes:
[0033] The pixel points are sequentially extracted from the first smoothed image, and the pixel gradient amplitude and gradient angle of the extracted pixel points are calculated, the positive and negative gradient intervals corresponding to the extracted pixel points are obtained based on the pixel gradient amplitude and gradient angle, the gray value extreme value in the positive and negative gradient intervals is identified, and the pixel point corresponding to the gray value extreme value is confirmed as a center pixel point, the center pixel points are summarized to obtain a center pixel point set, and an initial first edge image is constructed based on the center pixel point set;
[0034] The first smoothed image is subjected to Laplace transformation by using a pre-constructed Laplace formula, and the first smoothed image after Laplace transformation is confirmed as an initial second edge image, wherein the calculation formula of the Laplace transformation is as follows:
[0035]
[0036] wherein L(x i , y i ) represents a pixel value corresponding to a pixel coordinate in the initial second edge image, and f(x, y) represents a target text image.
[0037] The initial first edge image and the initial second edge image are fused based on a preset first enhancement weight and a preset second enhancement weight to obtain a text edge enhancement image.
[0038] Optionally, the optimized image is obtained based on the text edge enhancement image, comprising:
[0039] An initial neural network model is obtained, wherein the initial neural network model comprises a convolution layer, a maximum pooling layer and a global average pooling layer;
[0040] The convolution layer is optimized by using a pre-constructed batch normalization algorithm to obtain a suboptimal convolution layer, the text edge enhancement image is imported into the suboptimal convolution layer to obtain a suboptimal text image, the maximum pooling layer is used to perform a pooling operation on the suboptimal text image to obtain a first optimized image;
[0041] The first optimized image is subjected to a convolution operation by using a pre-constructed second suboptimal convolution layer to obtain a second optimized image, the number of iterations is counted based on the second optimized image, the second optimized image is taken as the text edge enhancement image, and the step of importing the text edge enhancement image into the suboptimal convolution layer to obtain the suboptimal text image is returned until the number of iterations is equal to a preset iteration threshold value, and a suboptimal output image is obtained;
[0042] A feature compression layer is constructed by using the global average pooling layer and the convolution layer, and the suboptimal output image is input into the feature compression layer, and the optimized image is obtained by using the suboptimal output image in the feature compression layer.
[0043] Optionally, the generating the text-enhanced image according to the optimized image and the pre-built decoder comprises:
[0044] performing a background removal operation on the optimized image by using a pre-built decoder to obtain the text-enhanced image, wherein the decoder is composed of a neural network, and the decoder comprises a decoding input layer, a first network layer, a second network layer and an output layer, wherein the first network layer is composed of a decoding suboptimal convolution layer and an up-sampling layer, and the second network layer is composed of a decoding suboptimal convolution layer.
[0045] Optionally, the constructing a unit time axis corresponding to the optimized text based on the set of time nodes comprises:
[0046] If there are multiple time nodes in the set of time nodes, performing a pairwise combination operation on the multiple time nodes to obtain a set of time node pairs, wherein the time node pair comprises a first time node and a second time node, sequentially extracting time node pairs from the set of time node pairs, and performing the following operations on the extracted time node pairs:
[0047] sorting the first time node and the second time node in the order from early to late according to time, and identifying the sorted first time node and the second time node by using a pre-built time completion structure to obtain a sequential time node pair;
[0048] constructing a first sequential time node and a second sequential time node based on the sequential time node pair;
[0049] summarizing the first sequential time node and the second sequential time node, and performing a deduplication operation on the summarized first sequential time node and the second sequential time node to obtain a unit time axis corresponding to the optimized text.
[0050] Optionally, the constructing a time-series evidence graph based on the evidence chain time axis and the set of optimized texts comprises:
[0051] performing an entity extraction operation on the optimized text by using a pre-built large language model to obtain an entity set, performing a pairwise combination operation on the entity set to obtain an entity pair set, sequentially extracting entity pairs from the entity pair set, and performing the following operations on the extracted entity pairs:
[0052] constructing a time-series entity pair by using a pre-built triple relationship structure and the entity pair, wherein the time-series entity pair comprises a first entity-time-second entity;
[0053] if the time sequence in the time-series entity pair is a preset unknown time sequence, identifying the time-series entity pair as an unordered entity pair;
[0054] Otherwise, a first accurate time corresponding to the first entity and a second accurate time corresponding to the second entity are identified respectively, if the first accurate time and the second accurate time are both preset empty time, the time sequence entity pair is identified as a non-reference entity pair, otherwise, the first accurate time and the second accurate time are used to identify the time sequence entity pair as a reference entity pair;
[0055] The reference entity pairs are summarized to obtain a reference entity pair set, and the unordered entity pairs and the non-reference entity pairs are summarized to obtain a to-be-analyzed entity pair set;
[0056] The reference entity pair set is time-decomposed using the large language model and the evidence chain timeline to obtain an initial time sequence evidence graph;
[0057] The to-be-analyzed entity pair set is sent to a pre-constructed user terminal, the received to-be-analyzed entity pair set is analyzed based on the user terminal, and the analyzed to-be-analyzed entity pair set is taken as the entity pair set, and the step of returning the time sequence in the time sequence entity pair being preset unknown time sequence is returned until the to-be-analyzed entity pair set is empty, and a time sequence evidence graph is obtained.
[0058] To achieve the above purpose, the present application also provides an evidence screening system based on a large language model, comprising:
[0059] An instruction receiving module is configured to receive an initial text evidence set, an initial image evidence set, and an initial audio evidence set;
[0060] A multi-modal processing module is configured to obtain a to-be-processed text set based on the initial text evidence set, the initial image evidence set, and the initial audio evidence set, wherein the to-be-processed text set includes a text evidence set, an image text set, and an audio text set, to-be-processed text sets are sequentially extracted from the to-be-processed text set, and an optimized text set is obtained by correcting the to-be-processed text set using a pre-constructed natural language model; an identified case data set is obtained using the historical case data set, and an initial retrieval database is constructed based on the identified case data set;
[0061] An evidence chain timeline processing module is configured to sequentially extract optimized texts from the optimized text set, and perform the following operations on the optimized texts: identifying a time node set corresponding to the optimized texts, constructing a unit timeline corresponding to the optimized texts based on the time node set, wherein the time node set includes zero, one, or more time nodes, summarizing the unit timelines to obtain a unit timeline set corresponding to the optimized text set, and constructing an evidence chain timeline based on the unit timeline set, wherein the evidence chain timeline includes a plurality of evidence time nodes;
[0062] A time sequence evidence graph construction module is configured to construct a time sequence evidence graph based on the evidence chain timeline and the optimized text set, and complete evidence screening based on the large language model.
[0063] To solve the above problems, the present application also provides an electronic device, which comprises:
[0064] A memory stores at least one instruction; and a processor executes the instruction stored in the memory to realize the evidence screening method based on a large language model.
[0065] To solve the above problems, the present application also provides a computer readable storage medium, which stores at least one instruction, and the at least one instruction is executed by a processor in an electronic device to realize the evidence screening method based on a large language model.
[0066] Compared with the prior art, the present application has the following beneficial effects:
[0067] The present application constructs a time sequence evidence graph based on an evidence chain timeline and an optimized text set, uses a large language model and an evidence chain timeline to time deduce a benchmark entity pair set, so that the benchmark entity set can be integrated into an initial suboptimal time sequence evidence graph composed of multiple nodes and corresponding benchmark entities in the order of occurrence of the benchmark entities, and further fuses evidence into the initial suboptimal time sequence evidence graph to obtain a time sequence evidence graph. Therefore, the present application can realize parsing and arrangement of evidence of different modalities in chronological order to assist in combing evidence. BRIEF DESCRIPTION OF DRAWINGS
[0068] Figure 1 A flowchart of the evidence screening method based on a large language model provided by the present application is shown.
[0069] Figure 2 A functional module diagram of the evidence screening system based on a large language model provided by the present application is shown.
[0070] Figure 3 A structural diagram of the electronic device for realizing the evidence screening method based on a large language model provided by the present application is shown.
[0071] REFERENCE SIGNS:
[0072] 1, electronic device; 10, processor; 11, memory; 12, bus.
[0073] The implementation, functional features and advantages of the present application will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0074] It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0075] Embodiment 1:
[0076] The embodiment provides a large language model-based evidence screening method, and an execution subject of the large language model-based evidence screening method includes but is not limited to at least one of electronic devices such as a server, a terminal and the like that can be configured to execute the method provided in the embodiment. In other words, the large language model-based evidence screening method can be executed by software or hardware installed in a terminal device or a server device, and the software can be a blockchain platform. The server includes but is not limited to a single server, a server cluster, a cloud server or a cloud server cluster and the like.
[0077] Referring to Figure 1 FIG. 1 is a flowchart of a large language model-based evidence screening method provided by an embodiment of the present application, and the large language model-based evidence screening method includes the following steps.
[0078] S1, receiving an initial text evidence set, an initial image evidence set and an initial audio evidence set.
[0079] It can be understood that the initial text evidence set is a collection of evidence recorded in the form of text in a case, the initial image evidence set is a collection of evidence recorded in the form of image in a case, and the initial audio evidence set is a collection of evidence recorded in the form of audio in a case.
[0080] For example, Xiao Zhang describes the process of event A, and the interrogator records Xiao Zhang's description, and Xiao Zhao also describes the process of event A, and the interrogator records Xiao Zhang's description. The interrogator also records the audio of Xiao Zhang and the audio of Xiao Zhao, respectively, and the collection of the content recorded by the interrogator and the content recorded by the interrogator constitutes the initial text evidence set, and the collection of the pictures taken at different positions in the scene of event A constitutes the initial audio evidence set.
[0081] It should be noted that in the normal narration of the process of event occurrence, people do not directly narrate the occurrence of events in chronological order, and in order to determine whether there is a lying behavior in the narration process, the time sequence is also disturbed to ask the narrator and answer what happened in a certain period of time, so in this case the initial data obtained is not recorded in chronological order, for example, Wang's narration of the events that occurred a few days ago is: "the event occurred yesterday or the day before yesterday, oh I came home yesterday, so the event occurred, today is February 25, so I went to A hotel the day before yesterday, and I booked a bouquet of flowers before going to A hotel", therefore, it is necessary to integrate these time sequence disturbed events based on time, so as to obtain an evidence chain in chronological order to help the inquiring personnel to sort out the whole process of the event occurrence.
[0082] S2, based on the initial text evidence set, the initial image evidence set and the initial audio evidence set, obtaining a to-be-processed text set, wherein the to-be-processed text set includes a text evidence set, an image text set and an audio text set.
[0083] Further, the to-be-processed text set is obtained based on the initial text evidence set, the initial image evidence set and the initial audio evidence set, including:
[0084] The initial image evidence is extracted from the initial image evidence set in sequence to obtain a target processing image, the image identification of the target processing image is generated based on the target processing image and the pre-constructed image identification structure, the identification operation is performed on the target processing image by using the image identification, the identified target processing image is summarized to obtain an identification image set, and the image text set is obtained based on the identification image set, wherein the image text set includes a plurality of image texts, and the image text corresponds to the image identification of the identification image one by one.
[0085] The initial audio evidence set is converted into a speech text set by using the pre-constructed speech identification structure, wherein the speech text set includes a plurality of speech texts, and the speech text corresponds to the speech identification one by one.
[0086] The speech text set, the image text set and the initial text evidence set are summarized to obtain the to-be-processed text set.
[0087] It should be noted that when analyzing an event, it should be analyzed from multiple angles and multiple aspects, therefore the embodiment of the present application receives an initial text evidence set, an initial image evidence set and an initial audio evidence set, but although the large language model has the ability to deduce text, it is essentially a text processing tool, and therefore cannot directly import pictures and audio into the large language model to analyze the whole process of event development, so the initial text evidence set, the initial image evidence set and the initial audio evidence set need to be processed, and all converted into a text form of a to-be-processed text set, and then the large language model is used to deduce the process of the event.
[0088] Further, the icon identification structure is a structure for identifying the target processing image, and the image identification of the target processing image based on the target processing image and the pre-constructed icon identification structure specifically refers to a method capable of integrating some information in the target processing image acquisition process into a piece of data in the embodiment of the present application, for example, an inquiring person makes a text identification on a target processing image, and the identified content and order are: the photographer, the shooting time, the shooting location, the event, the shooting content and the natural scene text, and the text identification is integrated into an XML format for storage, and therefore the image identification is the XML format text identification.
[0089] Further, the identification operation of the target processing image by using the image identification is to associate the image identification with the target processing image together, for example, the image identification is further simplified, and the simplified image identification is used as a key, and the target processing image is used as a value to construct a key-value pair to perform the identification operation, and for another example, the image identification and the target processing image can be simply stored as the same data.
[0090] Among them, the image text set includes a plurality of image texts, and the image text is a text composed of the shooting location, the event, the shooting content and the natural scene text extracted from the image identification.
[0091] At the same time, the voice identification structure is the same as the icon identification structure, the voice identification is the same as the image identification, and has similar effects, which will not be described here again. The voice text set is a set composed of text contents extracted from the initial audio evidence set. Since the initial audio evidence includes a plurality of initial audio evidences, the content of each initial audio evidence is of different length, and if the text of the initial audio is very long, it will greatly increase the workload of the computer, therefore, the initial audio evidence set is converted into a voice text set by using the pre-constructed voice identification structure, which includes:
[0092] The initial audio evidences are sequentially extracted from the initial audio evidence set, and the following operations are performed on the initial audio evidences:
[0093] The initial audio evidence is slidingly intercepted by using a pre-constructed sliding interception time length and a sliding interception step length, to obtain a plurality of unit initial audios, and texts corresponding to the plurality of unit initial audios are recognized to obtain a unit text set;
[0094] The unit text set is summarized to obtain a speech text set.
[0095] It should be noted that the sliding interception time length is similar to the sliding window, and the initial audio evidence corresponding audio can be slidingly intercepted by using a window with a preset time length; the interception step length is similar to the sliding step length corresponding to the sliding window concept, and is used to control the size of the sliding interception time length gradually moving.
[0096] Further, in the unit text set, the unit initial audio and the unit text are one-to-one corresponding, and the recognition method has a plurality of existing speech-to-text technologies that can be implemented, which will not be described here.
[0097] Further, the image identification of the target processing image is generated based on the target processing image and the pre-constructed icon identification structure, the identification operation is performed on the target processing image by using the image identification, the identified target processing image is summarized to obtain an identified image set, which includes:
[0098] If the target processing image does not exist in the preset text area, the target processing image is confirmed as a non-text image, and a non-text icon identification is generated based on the non-text image and the icon identification structure, and an identification operation is performed on the non-text image by using the non-text icon identification to obtain an identified image;
[0099] Otherwise, the target processing image is cropped based on the text area to obtain a target text image set, wherein the target text image set includes one or more target text images;
[0100] The following operations are performed on one or more target text images:
[0101] A convolution operation is performed on the target text image by using a pre-constructed Gaussian filter to obtain a first smoothed image, an edge enhancement operation is performed on the first smoothed image to obtain a text edge enhancement image, a background removal operation is performed on the text edge enhancement image to obtain a text enhancement image, a text recognition operation is performed based on the text enhancement image to obtain a target image text;
[0102] The target image text is summarized to obtain a target image text set, a text icon identification is constructed based on the target image text set and the icon identification structure, and an identification operation is performed on the target processing image by using the text icon identification to obtain an identified image;
[0103] The identified images are summarized to obtain an identified image set.
[0104] The text region is a text in the target processing image. For example, the target processing image is a signboard of a store, and the region corresponding to the signboard of the store is the text region. For another example, if the target processing image is a large tree, the target processing image does not include a text region. The no-text image is an image in the target processing image that does not include a text. It should be noted that the image identifier includes a no-text identifier and a text identifier. The no-text identifier is an identifier generated based on the image identifier structure and used to identify an image in the target processing image that does not include a text. For example, a person inquiring about the target processing image performs text identification, and the content and order of the identification are as follows: a person taking a picture, a picture taking time, a picture taking place, an event occurring, a picture taking content, and a natural scene text. Since the target processing image does not include any image in a natural scene, the natural scene text is NULL. The text identification is integrated into an XML format and stored in the picture taking content. At this time, the text identification is a no-text identifier.
[0105] Further, the target text image set is a set of images corresponding to all text regions in the target processing image. For example, the target processing image includes signboards of store A, store B, and store C. The text regions corresponding to the signboards of store A, store B, and store C are cropped and collected, and the obtained set is the target text image set.
[0106] It should be noted that the text identification operation is an operation of identifying a text from an image. For example, the target text is a text identified from a text enhanced image by using an OCR technology. The text identifier is similar to the no-text identifier and can achieve the same effect, but the difference is that the text identifier stores the text in the natural scene in the target text image.
[0107] It can be understood that the convolution operation is performed on the target text image by using a pre-constructed Gaussian filter to obtain a first smoothed image, an edge enhancement operation is performed on the first smoothed image to obtain a text edge enhanced image, and a background removal operation is performed on the text edge enhanced image to obtain a text enhanced image. The convolution operation includes the following steps.
[0108] A Gaussian filter is generated by using a pre-constructed two-dimensional Gaussian kernel, and a convolution operation is performed on the target text image by using the Gaussian filter to obtain a first smoothed image. The convolution formula is as follows.
[0109]
[0110] P=G*M
[0111] Wherein, x represents the horizontal coordinate of the target text image corresponding to the image coordinate system, y represents the vertical coordinate of the target text image corresponding to the image coordinate system, G(x, y) represents a two-dimensional Gaussian kernel, G represents a Gaussian filter constructed based on a Gaussian filter function, P represents the first smoothed image, and M represents the target text image.
[0112] Based on the first smoothed image and the pre-constructed edge enhancement method, a text edge enhancement image is obtained.
[0113] Based on the text edge enhancement image, an optimized image is obtained, and a text enhancement image is generated according to the optimized image and a pre-constructed decoder.
[0114] Therefore, in the embodiment, the Gaussian filter is used to perform convolution operation on the target text image to obtain the first smoothed image, which aims to eliminate the noise in the target text image by using the Gaussian filter, thereby reducing the irrelevant edges in the edge detection result caused by the existence of noise, which are not related to the text to be recognized, thereby reducing the calculation amount of recognizing the text from the picture, and improving the accuracy of text recognition.
[0115] Further, the text edge enhancement image is obtained based on the first smoothed image and the pre-constructed edge enhancement method, which includes:
[0116] The pixel points are extracted from the first smoothed image in sequence, and the pixel gradient amplitude and gradient angle of the extracted pixel points are calculated, the positive and negative gradient intervals corresponding to the extracted pixel points are obtained based on the pixel gradient amplitude and gradient angle, the gray value extreme value in the positive and negative gradient intervals is identified, and the pixel point corresponding to the gray value extreme value is confirmed as the center pixel point. The center pixel points are collected to obtain a center pixel point set, and an initial first edge image is constructed based on the center pixel point set.
[0117] The Laplace transform of the first smoothed image is performed by using the pre-constructed Laplace formula, and the first smoothed image after the Laplace transform is confirmed as the initial second edge image, wherein the calculation formula of the Laplace transform is as follows:
[0118]
[0119] Wherein, L(x i , y i ) represents the pixel value corresponding to the pixel coordinate in the initial second edge image, and f(x, y) represents the target text image.
[0120] The initial first edge image and the initial second edge image are fused based on the pre-set first enhancement weight and the pre-set second enhancement weight to obtain the text edge enhancement image.
[0121] It can be understood that the calculation formula of the pixel gradient amplitude and the gradient angle is as follows:
[0122]
[0123] wherein, t p represents the component of the gradient of the pixel point (p, q) in the first smoothed image in the horizontal axis direction, t q represents the component of the gradient of the pixel point (p, q) in the first smoothed image in the vertical axis direction, p represents the horizontal axis in the image coordinate system corresponding to the first smoothed image, q represents the vertical axis in the image coordinate system corresponding to the first smoothed image, P(p+1, q) represents the gray value corresponding to the pixel point with coordinates (p+1, q) in the first smoothed image, P(p-1, q) represents the gray value corresponding to the pixel point with coordinates (p-1, q) in the first smoothed image, P(p, q+1) represents the gray value corresponding to the pixel point with coordinates (p, q+1) in the first smoothed image, P(p, q-1) represents the gray value corresponding to the pixel point with coordinates (p, q-1) in the first smoothed image, t θ represents the gradient amplitude, and θ represents the gradient angle.
[0124] Further, the positive and negative gradient interval is an interval formed by the gradient amplitudes of the other multiple pixel points in the positive gradient direction and the negative gradient direction of the gradient angle of the extracted pixel point, and the number of the other multiple pixel points can be artificially set. The gray value extreme value is the maximum value of all the gray values in the positive and negative gradient interval, and the method of constructing the initial first edge image based on the center pixel point set can be the method of constructing the image edge in the canny edge detection algorithm.
[0125] It should be noted that the first enhancement weight is the weight of the initial first edge image in the fusion process, the second enhancement weight is the weight of the initial second edge image in the fusion process, and the character edge enhanced image is an image obtained by weighting and summing the initial first edge image and the initial second edge image after using the first enhancement weight and the second enhancement weight.
[0126] Further, the method of obtaining the optimized image based on the character edge enhanced image comprises:
[0127] obtaining an initial neural network model, wherein the initial neural network model comprises a convolution layer, a maximum pooling layer and a global average pooling layer;
[0128] optimizing the convolution layer by using a pre-constructed batch normalization algorithm to obtain a sub-optimal convolution layer, importing the character edge enhanced image into the sub-optimal convolution layer to obtain a sub-optimal character image, performing a pooling operation on the sub-optimal character image by using the maximum pooling layer to obtain a first optimized image;
[0129] performing a convolution operation on the first optimized image by using a second optimized convolution layer pre-built, obtaining a second optimized image, counting an iteration number based on the second optimized image, taking the second optimized image as the text edge enhancement image, returning the step of importing the text edge enhancement image into the sub-optimal convolution layer to obtain a sub-optimal text image until the iteration number is equal to a preset iteration threshold, and obtaining a sub-optimal output image;
[0130] building a feature compression layer by using a global average pooling layer and the convolution layer, inputting the sub-optimal output image into the feature compression layer, and obtaining an optimized image by using the sub-optimal output image in the feature compression layer.
[0131] It should be noted that the method of optimizing the convolution layer by using a pre-built batch normalization algorithm to obtain a sub-optimal convolution layer is that in the convolution layer, the parameters input into the convolution layer are normalized to a standard normal distribution with a mean of 0 and a variance of 1, thereby improving the training speed and the ability to process images. The max pooling layer is a neural network layer that processes the sub-optimal text image by using a max pooling algorithm. The second optimized convolution layer has the same structure as the sub-optimal convolution layer and can achieve the same effect. The iteration number is the number of times of obtaining the second optimized image, for example, when the second optimized image is obtained for the first time, the iteration number is 1. The iteration threshold is the maximum number of iteration numbers set by a person. The feature compression layer is a neural network layer that processes the sub-optimal text image by using an average pooling algorithm and a convolution layer. The optimized image is an image obtained by processing the sub-optimal output image by using the feature compression layer.
[0132] Further, the method of generating a text enhancement image according to the optimized image and a pre-built decoder comprises:
[0133] performing a background removal operation on the optimized image by using a pre-built decoder to obtain a text enhancement image, wherein the decoder is composed of a neural network, and the decoder comprises a decoding input layer, a first network layer, a second network layer and an output layer, wherein the first network layer is composed of a decoding sub-optimal convolution layer and an up-sampling layer, and the second network layer is composed of a decoding sub-optimal convolution layer.
[0134] It should be noted that the background removal operation is an operation of removing the background part of the optimized image. The decoding input layer is the input layer of the decoder, and the decoding sub-optimal convolution layer is a network node in the decoder with the same structure as the sub-optimal convolution layer, so the first network layer is composed of the decoding sub-optimal convolution layer and the up-sampling layer.
[0135] S3, sequentially extracting the to-be-processed text set from the to-be-processed text set, and correcting the to-be-processed text set by using the pre-constructed natural language model to obtain an optimized text set. In this step, the natural language model can identify the incorrect content in each to-be-processed text in the to-be-processed text set due to wrong words or inaccurate recognition, and correct the incorrect content. Therefore, the optimized text set is the corrected to-be-processed text set.
[0136] S4, sequentially extracting the optimized text from the optimized text set, and performing the following operations on the optimized text: identifying a time node set corresponding to the optimized text, and constructing a unit time axis corresponding to the optimized text based on the time node set, wherein the time node set includes zero, one or more time nodes.
[0137] It should be noted that the time node set is a set of words with explicit time meaning extracted from the optimized text, which can be identified by using a large language model based on a pre-constructed corpus. For example, Xiaowang describes the events that occurred a few days ago as follows: "The event occurred yesterday or the day before yesterday, oh I went home yesterday, so the event occurred, today is February 25, so I went to A hotel the night before yesterday, and I booked a bouquet of flowers before going to A hotel". From this, we can extract
yesterday
the day before yesterday
February 25
the night before yesterday
[0138] Further, the construction of the unit time axis corresponding to the optimized text based on the time node set comprises:
[0139] If there are multiple time nodes in the time node set, perform a two-by-two combination operation on the multiple time nodes to obtain a time node pair set, wherein the time node pair includes a first time node and a second time node, sequentially extract the time node pairs from the time node pair set, and perform the following operations on the extracted time node pairs:
[0140] Sort the first time node and the second time node in the order of time from early to late, and identify the sorted first time node and the second time node by using the pre-constructed time completion structure to obtain a sequential time node pair;
[0141] Constructing a first sequential time node and a second sequential time node based on the sequential time node pair;
[0142] Summarizing the first sequential time node and the second sequential time node, and performing a deduplication operation on the summarized first sequential time node and the second sequential time node to obtain a unit time axis corresponding to the optimized text.
[0143] It can be understood that the pairwise combination operation is an operation of combining multiple time nodes, for example, the multiple time nodes are
yesterday
the day before yesterday
February 25th
the day before yesterday evening
yesterday, the day before yesterday
yesterday, February 25th
yesterday, the day before yesterday evening
the day before yesterday, February 25th
the day before yesterday, the day before yesterday evening
February 25th, the day before yesterday evening
yesterday, the day before yesterday evening
yesterday
the day before yesterday evening
[0144] Further, when the first time node and the second time node are both sorted in the order from early to late, the large language model can understand the order between the first time node and the second time node according to the semantics of the received optimization text. Optionally, the time completion structure can be in the ISO 8601 format and adjusted according to human intent. For example, during the process of an event, the accuracy only needs to be accurate to the morning, noon, afternoon, evening, etc., and the ISO 8601 format can be appropriately adjusted to meet the actual use requirements. The method of identifying the sorted first time node and the second time node using the time completion structure can use the index marking method for marking.
[0145] Further, the first order time node and the second order time node are two fixed nodes constructed according to the order time node pair. The execution of the deduplication operation is to remove the first order time node or the second order time node that is repeated in time in the aggregated first order time node and the second order time node.
[0146] For example, there are two time node pairs
Yesterday, the day before yesterday
Yesterday, February 25th
Yesterday, February 25th
Yesterday, the day before yesterday
Yesterday, the day before yesterday
The day before yesterday (February 23rd), yesterday (February 24th)
Yesterday, February 25th
Yesterday (February 24th), February 25th
The day before yesterday (February 23rd), yesterday (February 24th)
Yesterday (February 24th), February 25th
The day before yesterday (February 23rd)-yesterday (February 24th)-February 25th
[0147] S5, aggregating the unit time axis to obtain a unit time axis set corresponding to the optimization text set.
[0148] It should be noted that the unit time axis corresponding to different optimization texts is different, so the optimization text and the unit time axis correspond one by one.
[0149] S6, constructing an evidence chain time axis based on the unit time axis set, wherein the evidence chain time axis comprises a plurality of evidence time nodes.
[0150] Further, the unit time axis corresponding to different optimization texts is different, and there may be a part of the unit time axis that coincides. Therefore, when constructing the evidence chain time axis, the repeated nodes in the unit time axis are de-duplicated to ensure that the nodes in the evidence chain time axis are unique.
[0151] S7, constructing a time sequence evidence graph based on the evidence chain time axis and the optimization text set, and completing the evidence screening based on the large language model.
[0152] Further, the constructing a time sequence evidence graph based on the evidence chain time axis and the optimization text set comprises:
[0153] The pre-constructed large language model is used for entity extraction operation on the optimized text to obtain an entity set, a two-by-two combination operation is performed on the entity set to obtain an entity pair set, an entity pair is extracted from the entity pair set in sequence, and the following operations are performed on the extracted entity pair:
[0154] A pre-constructed triple relationship structure and an entity pair are used to construct a time sequence entity pair, wherein the time sequence entity pair comprises a first entity-time sequence-second entity.
[0155] If the time sequence in the time sequence entity pair is a preset unknown time sequence, the time sequence entity pair is identified as an unordered entity pair.
[0156] Otherwise, a first accurate time corresponding to the first entity and a second accurate time corresponding to the second entity are identified, if the first accurate time and the second accurate time are both preset empty time, the time sequence entity pair is identified as a no-reference entity pair, otherwise, the time sequence entity pair is identified as a reference entity pair using the first accurate time and the second accurate time.
[0157] The reference entity pairs are summarized to obtain a reference entity pair set, and the unordered entity pairs and the no-reference entity pairs are summarized to obtain a to-be-analyzed entity pair set.
[0158] The large language model and the evidence chain timeline are used to perform time deduction on the reference entity pair set to obtain an initial time sequence evidence graph.
[0159] The to-be-analyzed entity pair set is sent to a pre-constructed user terminal, the received to-be-analyzed entity pair set is analyzed based on the user terminal, and the analyzed to-be-analyzed entity pair set is taken as the entity pair set, the step of if the time sequence in the time sequence entity pair is a preset unknown time sequence is returned, until the to-be-analyzed entity pair set is an empty set, and a time sequence evidence graph is obtained.
[0160] It should be noted that the method of performing a two-by-two combination operation on the entity set to obtain an entity pair set is the same as the method of performing a two-by-two combination operation on a plurality of time nodes to obtain a time node pair set, which will not be repeated here. The triple relationship structure is a structure composed of a first entity, a time sequence, and a second entity, and the first entity and the second entity are two entities in an entity pair. The time sequence in the time sequence entity pair includes "before" and "at the same time". The unknown time sequence is an entity pair in which it is unknown whether the first entity precedes the second entity, the second entity precedes the first entity, or the first entity and the second entity occur at the same time, that is, the unknown time sequence indicates that the time sequence between the first entity and the second entity cannot be identified, so the unordered entity pair indicates that the order between the first entity and the second entity in the entity pair is unknown.
[0161] The first precise time is the time when the first entity occurs, and the second precise time is the time when the second entity occurs. The first precise time and the second precise time are obtained by a method of understanding through optimizing the semantics of the text. The null time is the time when the first precise time or the second precise time cannot be identified.
[0162] The accuracy of the first precise time and the second precise time can be determined by a human being. Optionally, the accuracy in the embodiment of the present application is year + month + day + morning or afternoon. If the year + month + day + morning or afternoon can be identified, it is considered that the first precise time and the second precise time are accurate, and both are not null time. When the accuracy of the first precise time and the second precise time is lower than the preset accuracy, the first precise time and the second precise time are confirmed as null time. When the accuracy of the first precise time and the second precise time exceeds the preset accuracy, the most accurate time recording method is used for recording.
[0163] It should be further pointed out that as long as the first precise time and the second precise time are not both null time, the time sequence between the first entity and the second entity can be deduced by a large language model, and the first entity and the second entity can be positioned on the evidence chain time axis. For example, the evidence chain time axis is 2 / 23 morning-2 / 23 afternoon-2 / 25 morning, the precise time corresponding to the first entity is 2 / 23 afternoon, the first entity is earlier than the second entity, and therefore the possible occurrence range of the second entity is narrowed to 2 / 23 afternoon-2 / 25. In this way, the large language model and the evidence chain time axis can be used to time deduce the reference entity set, so that the reference entity set can be integrated into an initial suboptimal time sequence evidence graph which describes the reference entities in the order of occurrence of the reference entities and is composed of multiple nodes and corresponding reference entities.
[0164] Further, if the reference entity is traced, it is found that the reference entity comes from an initial image evidence, and the initial image evidence can be stored in the node corresponding to the reference entity in the initial suboptimal time sequence evidence graph to obtain an initial time sequence evidence graph.
[0165] Specifically, the user end is a port operated by an inquiry personnel or a human being. In the process of analyzing the set of entities to be analyzed by the user end, the user can analyze each entity pair in the set of entities to be analyzed, so as to analyze the sequence or occurrence of the two entities in the entity pair to be analyzed.
[0166] Thus, the embodiment can parse and arrange the evidence of different modalities in chronological order to assist in sorting out the evidence. Specifically, it first obtains a to-be-processed text set based on the initial text evidence set, the initial image evidence set, and the initial audio evidence set, converts all of them into a to-be-processed text set in text form, then deduces the process of the event occurrence through a large language model, simultaneously extracts the to-be-processed text set from the to-be-processed text set, and corrects the errors in the to-be-processed text set by using a pre-constructed natural language model, so as to make the time sequence evidence picture more accurate.
[0167] Further, the repeated nodes in the unit time axis are de-duplicated to ensure that the nodes in the evidence chain time axis are unique, thereby sorting out the chronological order of the occurrence time. The large language model and the evidence chain time axis are used to time deduce the benchmark entity pair set, so that the benchmark entity set can be integrated into an initial suboptimal time sequence evidence graph which is described in chronological order of the occurrence of the benchmark entity, and is composed of multiple nodes and corresponding benchmark entities, and the evidence is further fused into the initial suboptimal time sequence evidence graph to obtain a time sequence evidence graph, thereby greatly saving resources such as manpower and improving the efficiency of data sorting and collection.
[0168] Embodiment 2
[0169] As shown in Figure 2 The embodiment provides a large language model-based evidence screening system 100 which can be installed in an electronic device and includes an instruction receiving module 101, a multi-modal processing module 102, an evidence chain time axis processing module 103, and a time sequence evidence graph construction module 104. The modules of the present application can also be referred to as units, which are a series of computer program segments that can be executed by an electronic device processor and can complete a fixed function, and are stored in the memory of the electronic device.
[0170] The instruction receiving module 101 is configured to receive an initial text evidence set, an initial image evidence set, and an initial audio evidence set.
[0171] The multi-modal processing module 102 is configured to obtain a to-be-processed text set based on the initial text evidence set, the initial image evidence set, and the initial audio evidence set, wherein the to-be-processed text set includes a text evidence set, an image text set, and an audio text set, the to-be-processed text set is extracted from the to-be-processed text set, and the to-be-processed text set is corrected by using a pre-constructed natural language model to obtain an optimized text set.
[0172] The evidence chain timeline processing module 103 is configured to sequentially extract the optimization texts from the optimization text set, and perform the following operations on each optimization text: identifying a time node set corresponding to the optimization text, constructing a unit timeline corresponding to the optimization text based on the time node set, wherein the time node set includes zero, one or more time nodes, aggregating the unit timelines to obtain a unit timeline set corresponding to the optimization text set, and constructing an evidence chain timeline based on the unit timeline set, wherein the evidence chain timeline includes a plurality of evidence time nodes.
[0173] The time sequence evidence graph construction module 104 is configured to construct a time sequence evidence graph based on the evidence chain timeline and the optimization text set, and complete the evidence screening based on the large language model.
[0174] In detail, the modules in the evidence screening system based on the large language model 100 in the embodiments of the present application use the same technical means as the evidence screening method based on the large language model in the Figure 1 above and can produce the same technical effects, which will not be described here.
[0175] As shown in Figure 3 , it is a structural schematic diagram of an electronic device for implementing the evidence screening method based on the large language model according to an embodiment of the present application.
[0176] The electronic device 1 can include a processor 10, a memory 11 and a bus 12, and can further include a computer program stored in the memory 11 and executable on the processor 10, such as the evidence screening method based on the large language model program in Embodiment 1, and when the computer program is running in the processor 10, the evidence screening method based on the large language model in Embodiment 1 can be implemented.
[0177] Specifically, the specific implementation method of the processor 10 on the above instructions can refer to the description of the related steps in the corresponding embodiments, which will not be described here. Figures 1 to 3
[0178] Further, the modules / units integrated in the electronic device 1, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. The computer readable storage medium can be volatile or non-volatile. For example, the computer readable medium can include any entity or system capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory).
[0179] The application further provides a computer readable storage medium, the readable storage medium stores a computer program, and the computer program can implement the evidence screening method based on the large language model when being executed by a processor of an electronic device.
[0180] In several embodiments provided by the present application, it should be understood that the disclosed devices, systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative, and actual implementation can have other division ways.
[0181] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.
[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application, and although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A large language model-based evidence screening method, characterized in that, The method comprises: receiving an initial text evidence set, an initial image evidence set, and an initial audio evidence set; obtaining a to-be-processed text set based on the initial text evidence set, the initial image evidence set, and the initial audio evidence set, wherein the to-be-processed text set comprises a text evidence set, an image text set, and an audio text set; extracting the to-be-processed text set from the to-be-processed text set in turn, and correcting the to-be-processed text set by using a pre-constructed natural language model to obtain an optimized text set; extracting the optimized text from the optimized text set in turn, and performing the following operations on the optimized text: identifying a time node set corresponding to the optimized text, and constructing a unit time axis corresponding to the optimized text based on the time node set, wherein the time node set comprises zero, one, or more time nodes; aggregating the unit time axes to obtain a unit time axis set corresponding to the optimized text set, and constructing an evidence chain time axis based on the unit time axis set, wherein the evidence chain time axis comprises a plurality of evidence time nodes; constructing a time-series evidence graph based on the evidence chain time axis and the optimized text set to complete evidence screening based on a large language model, wherein the construction of the time-series evidence graph based on the evidence chain time axis and the optimized text set comprises: performing entity extraction on the optimized text by using a pre-constructed large language model to obtain an entity set, performing two-by-two combination operations on the entity set to obtain an entity pair set, extracting the entity pairs from the entity pair set in turn, and performing the following operations on the extracted entity pairs: constructing a time-series entity pair by using a pre-constructed triple relationship structure and the entity pairs, wherein the time-series entity pair comprises a first entity-time-second entity; if the time in the time-series entity pair is a preset unknown time, identifying the time-series entity pair as an unordered entity pair; otherwise, identifying a first accurate time corresponding to the first entity and a second accurate time corresponding to the second entity, respectively, if both the first accurate time and the second accurate time are a preset empty time, identifying the time-series entity pair as a no-reference entity pair, otherwise, identifying the time-series entity pair as a reference entity pair by using the first accurate time and the second accurate time; aggregating the reference entity pairs to obtain a reference entity pair set, and aggregating the unordered entity pairs and the no-reference entity pairs to obtain a to-be-analyzed entity pair set; performing time deduction on the reference entity pair set by using the large language model and the evidence chain time axis to obtain an initial time-series evidence graph; and sending the to-be-analyzed entity pair set to a pre-constructed user terminal, analyzing the received to-be-analyzed entity pair set based on the user terminal, and returning to the step of identifying the time-series entity pair as an unordered entity pair if the time in the time-series entity pair is a preset unknown time, until the to-be-analyzed entity pair set is empty, to obtain a time-series evidence graph.
2. The evidence screening method of claim 1, wherein, The obtaining of the to-be-processed text set based on the initial text evidence set, the initial image evidence set, and the initial audio evidence set comprises: Extracting initial image evidences from the initial image evidence set in sequence to obtain a target processing image, generating an image identifier of the target processing image based on the target processing image and a pre-constructed image identifier structure, performing an identification operation on the target processing image by using the image identifier, and collecting the identified target processing image to obtain an identified image set, and obtaining an image text set based on the identified image set, wherein the image text set includes a plurality of image texts, and the image texts correspond to the image identifiers of the identified images in a one-to-one manner; Converting the initial audio evidence set into a speech text set by using a pre-constructed speech identifier structure, wherein the speech text set includes a plurality of speech texts, and the speech texts correspond to the speech identifiers in a one-to-one manner; Collecting the speech text set, the image text set and the initial text evidence set to obtain a to-be-processed text set.
3. The evidence screening method of claim 2, wherein, The method further includes the following steps: If the target processing image does not include a preset text region, the target processing image is determined as a non-text image, a non-text image identifier is generated based on the non-text image and the image identifier structure, and an identification operation is performed on the non-text image by using the non-text image identifier to obtain an identified image; Otherwise, the target processing image is cropped based on the text region to obtain a target text image set, wherein the target text image set includes one or more target text images; The following operations are performed on each target text image in the one or more target text images: A first smoothing image is obtained by performing a convolution operation on the target text image by using a pre-constructed Gaussian filter, an edge enhancement operation is performed on the first smoothing image to obtain a text edge enhancement image, a background removal operation is performed on the text edge enhancement image to obtain a text enhancement image, a text recognition operation is performed on the text enhancement image to obtain a target image text, and the target image text is collected to obtain a target image text set, a text identifier is constructed based on the target image text set and the image identifier structure, and an identification operation is performed on the target processing image by using the text identifier to obtain an identified image; The identified images are collected to obtain an identified image set. The following operations are performed on each target text image in the one or more target text images:
4. The evidence screening method of claim 3, wherein, A first smoothing image is obtained by performing a convolution operation on the target text image by using a pre-constructed Gaussian filter, an edge enhancement operation is performed on the first smoothing image to obtain a text edge enhancement image, a background removal operation is performed on the text edge enhancement image to obtain a text enhancement image, a text recognition operation is performed on the text enhancement image to obtain a target image text, and the target image text is collected to obtain a target image text set, a text identifier is constructed based on the target image text set and the image identifier structure, and an identification operation is performed on the target processing image by using the text identifier to obtain an identified image; A Gaussian filter is generated by using a pre-constructed two-dimensional Gaussian kernel, and a convolution operation is performed on the target text image by using the Gaussian filter to obtain a first smoothed image, wherein the convolution formula is as shown in the following: wherein x represents a horizontal coordinate of an image coordinate system corresponding to the target text image, y represents a vertical coordinate of the image coordinate system corresponding to the target text image, G(x, y) represents a two-dimensional Gaussian kernel, G represents a Gaussian filter constructed based on a Gaussian filtering function, P represents the first smoothed image, and M represents the target text image. The identified images are collected to obtain an identified image set. The following operations are performed on each target text image in the one or more target text images:
5. The evidence screening method of claim 4, wherein, A first smoothing image is obtained by performing a convolution operation on the target text image by using a pre-constructed Gaussian filter, an edge enhancement operation is performed on the first smoothing image to obtain a text edge enhancement image, a background removal operation is performed on the text edge enhancement image to obtain a text enhancement image, a text recognition operation is performed on the text enhancement image to obtain a target image text, and the target image text is collected to obtain a target image text set, a text identifier is constructed based on the target image text set and the image identifier structure, and an identification operation is performed on the target processing image by using the text identifier to obtain an identified image; The pixel points are sequentially extracted from the first smoothed image, the pixel gradient amplitude and the gradient angle of the extracted pixel points are calculated, the positive and negative gradient intervals corresponding to the extracted pixel points are obtained based on the pixel gradient amplitude and the gradient angle, the gray value extreme value in the positive and negative gradient intervals is identified, and the pixel point corresponding to the gray value extreme value is confirmed as a center pixel point. The center pixel points are collected to obtain a center pixel point set, and an initial first edge image is constructed based on the center pixel point set; The first smoothed image is subjected to Laplace transformation using a pre-constructed Laplace formula, and the first smoothed image after the Laplace transformation is confirmed as an initial second edge image, wherein the calculation formula of the Laplace transformation is as follows: wherein, represents a pixel value corresponding to a pixel coordinate in the initial second edge image, represents a target text image; The initial first edge image and the initial second edge image are fused based on the preset first enhancement weight and the preset second enhancement weight to obtain a character edge enhancement image.
6. The evidence screening method of claim 5, wherein, The optimized image is obtained based on the character edge enhancement image, which includes: An initial neural network model is obtained, wherein the initial neural network model includes a convolution layer, a maximum pooling layer, and a global average pooling layer; The convolution layer is optimized using a pre-constructed batch normalization algorithm to obtain a sub-optimal convolution layer. The character edge enhancement image is imported into the sub-optimal convolution layer to obtain a sub-optimal character image. The maximum pooling layer is used to perform a pooling operation on the sub-optimal character image to obtain a first optimized image. A second sub-optimal convolution layer is used to perform a convolution operation on the first optimized image to obtain a second optimized image. The number of iterations is counted based on the second optimized image. The second optimized image is used as the character edge enhancement image. The step of importing the character edge enhancement image into the sub-optimal convolution layer to obtain a sub-optimal character image is returned until the number of iterations equals a preset iteration threshold to obtain a sub-optimal output image. A feature compression layer is constructed using the global average pooling layer and the convolution layer, and the sub-optimal output image is input into the feature compression layer. The optimized image is obtained using the sub-optimal output image in the feature compression layer.
7. The evidence screening method of claim 6, wherein, The text enhancement image is generated based on the optimized image and a pre-constructed decoder, which includes: A background removal operation is performed on the optimized image using a pre-constructed decoder to obtain a text enhancement image. The decoder is composed of a neural network, and the decoder includes a decoding input layer, a first network layer, a second network layer, and an output layer. The first network layer is composed of a decoding sub-optimal convolution layer and an up-sampling layer. The second network layer is composed of a decoding sub-optimal convolution layer. 8.The big language model-based evidence screening method of claim 7, wherein, The unit time axis corresponding to the optimized text is constructed based on the time node set, which includes: If there are multiple time nodes in the time node set, a pairwise combination operation is performed on the multiple time nodes to obtain a time node pair set. The time node pair includes a first time node and a second time node. The time node pairs are sequentially extracted from the time node pair set, and the following operations are performed on the extracted time node pairs: The first time node and the second time node are sorted in the order of time from early to late, and the sorted first time node and the second time node are identified using a pre-constructed time completion structure to obtain a sequential time node pair; The first sequential time node and the second sequential time node are constructed based on the sequential time node pair; The first sequential time node and the second sequential time node are collected, and a deduplication operation is performed on the collected first sequential time node and the second sequential time node to obtain a unit time axis corresponding to the optimized text. 9.A large language model-based evidence screening system, characterized in that, The system comprises: An instruction receiving module for receiving an initial text evidence set, an initial image evidence set and an initial audio evidence set; A multi-modal processing module for obtaining a to-be-processed text set based on the initial text evidence set, the initial image evidence set and the initial audio evidence set, wherein the to-be-processed text set comprises a text evidence set, an image text set and an audio text set, the to-be-processed text set is extracted from the to-be-processed text set in turn, and the to-be-processed text set is proofread by using a pre-constructed natural language model to obtain an optimized text set; An evidence chain timeline processing module for extracting the optimized text from the optimized text set in turn, and performing the following operations on the optimized text: identifying a time node set corresponding to the optimized text, constructing a unit timeline corresponding to the optimized text based on the time node set, wherein the time node set comprises zero, one or more time nodes, aggregating the unit timelines to obtain a unit timeline set corresponding to the optimized text set, and constructing an evidence chain timeline based on the unit timeline set, wherein the evidence chain timeline comprises a plurality of evidence time nodes; A time-series evidence graph construction module for constructing a time-series evidence graph based on the evidence chain timeline and the optimized text set, and completing evidence screening based on a large language model, wherein the construction of the time-series evidence graph based on the evidence chain timeline and the optimized text set comprises: performing entity extraction on the optimized text by using a pre-constructed large language model to obtain an entity set, performing two-by-two combination on the entity set to obtain an entity pair set, extracting the entity pairs from the entity pair set in turn, and performing the following operations on the extracted entity pairs: constructing a time-series entity pair by using a pre-constructed triple relationship structure and the entity pairs, wherein the time-series entity pair comprises a first entity-time-series-a second entity; if the time series in the time-series entity pair is a preset unknown time series, the time-series entity pair is identified as an unordered entity pair; otherwise, a first accurate time corresponding to the first entity and a second accurate time corresponding to the second entity are identified respectively, if the first accurate time and the second accurate time are both preset empty time, the time-series entity pair is identified as a no-reference entity pair, otherwise, the time-series entity pair is identified as a reference entity pair by using the first accurate time and the second accurate time; aggregating the reference entity pairs to obtain a reference entity pair set, and aggregating the unordered entity pairs and the no-reference entity pairs to obtain a to-be-analyzed entity pair set; performing time deduction on the reference entity pair set by using the large language model and the evidence chain timeline to obtain an initial time-series evidence graph; and sending the to-be-analyzed entity pair set to a pre-constructed user terminal, analyzing the received to-be-analyzed entity pair set based on the user terminal, and returning to the step of if the time series in the time-series entity pair is a preset unknown time series until the to-be-analyzed entity pair set is empty to obtain the time-series evidence graph.
Citation Information
Patent Citations
Method for achieving electronic evidence data analysis based on time tracks
CN103729397A
Method for associating multiple evidences in case records based on knowledge graph
CN118761475A