如何删除java中的停用词?(How to remove stop words in java?)
我想删除java中的停用词。
所以,我从文本文件中读取了停用词。
并存储设置
Set<String> stopWords = new LinkedHashSet<String>(); BufferedReader br = new BufferedReader(new FileReader("stopwords.txt")); String words = null; while( (words = br.readLine()) != null) { stopWords.add(words.trim()); } br.close();而且,我读了另一个文本文件。
所以,我想删除文本文件中的重复字符串。
我怎么能够?
I want to remove stop words in java.
So, I read stop words from text file.
and store Set
Set<String> stopWords = new LinkedHashSet<String>(); BufferedReader br = new BufferedReader(new FileReader("stopwords.txt")); String words = null; while( (words = br.readLine()) != null) { stopWords.add(words.trim()); } br.close();And, I read another text file.
So, I wanna remove to duplicate string in text file.
How can I?
最满意答案
你想从文件中删除重复的单词,下面是相同的高级逻辑。
读取文件 循环播放文件内容(即一次一行) 根据空间为该行提供字符串标记生成器 将每个标记添加到您的设置中。 这将确保您每个单词只有一个条目。 关闭文件现在你已经设置了包含文件的所有唯一字。
You want to remove duplicate words from file, below is the high level logic for same.
Read File Loop through file content(i.e one line at a time) Have string tokenizer for that line based on space Add each each token to your set. This will make sure that you have only one entry per word. Close fileNow you have set that contains all the unique word of file.
更多推荐
发布评论