Skip to content

Html Content / Article Extractor in Golang

License

Notifications You must be signed in to change notification settings

shaneiseminger/GoOse

This branch is 13 commits behind advancedlogic/GoOse:master.

Folders and files

NameName
Last commit message
Last commit date

Latest commit

3b3ad76 · Aug 22, 2020
Jan 20, 2017
Nov 4, 2015
Nov 22, 2015
Aug 25, 2018
Aug 25, 2018
Jan 10, 2014
May 1, 2017
May 18, 2017
May 1, 2017
Oct 12, 2018
Jun 6, 2017
Dec 1, 2015
Dec 4, 2017
Mar 18, 2018
Nov 24, 2015
Nov 12, 2019
Oct 2, 2019
Oct 7, 2019
May 1, 2017
Nov 12, 2019
Nov 12, 2019
Oct 28, 2019
Oct 4, 2019
Oct 19, 2019
Jan 11, 2014
Aug 22, 2020
Feb 18, 2019
May 1, 2017
Jan 26, 2017
Aug 25, 2018
Aug 25, 2018
Aug 25, 2018

Repository files navigation

GoOse

HTML Content / Article Extractor in Golang

Build Status Coverage Status Go Report Card GoDoc

Description

This is a golang port of "Goose" originaly licensed to Gravity.com under one or more contributor license agreements. See the NOTICE file distributed with this work for additional information regarding copyright ownership.

Golang port was written by Antonio Linari

Gravity.com licenses this file to you under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

INSTALL

go get github.com/advancedlogic/GoOse

HOW TO USE IT

package main

import (
	"github.com/advancedlogic/GoOse"
)

func main() {
	g := goose.New()
	article, _ := g.ExtractFromURL("http://edition.cnn.com/2012/07/08/opinion/banzi-ted-open-source/index.html")
	println("title", article.Title)
	println("description", article.MetaDescription)
	println("keywords", article.MetaKeywords)
	println("content", article.CleanedText)
	println("url", article.FinalURL)
	println("top image", article.TopImage)
}

Development - Getting started

This application is written in GO language, please refere to the guides in https://golang.org for getting started.

This project include a Makefile that allows you to test and build the project with simple commands. To see all available options:

make help

Before committing the code, please check if it passes all tests using

make deps
make qa

TODO

  • better organize code
  • improve "xpath" like queries
  • add other image extractions techniques (imagemagick)

THANKS TO

@Martin Angers for goquery
@Fatih Arslan for set
GoLang team for the amazing language and net/html

About

Html Content / Article Extractor in Golang

Resources

License

Stars

Watchers

Forks

Releases

No releases published

Packages

No packages published

Languages

  • HTML 95.8%
  • Go 4.1%
  • Other 0.1%