Determine active nodes on network

Started by overdrive, January 15, 2014, 02:59:44 PM

overdrive

Hi
First off, thank you for an awesome library and good ideas.

I have changed out the rf portion of all my projects to use the rfm69hw and the lowpowerlab library. I have various custom "arduinos" running around the place (10 for now but soon it can be 20+). I am looking for a good/reliable way to find out "who" is out there and active on my rf network segment. I have a gateway node connected to usb via ftdi running a atmega328p-au. 1:1 communication to any node is reliable and fast.
In order to find out what nodes are active and their type I have tried setting the modules in promiscuous mode and have the gateway send out a "WHO" command and listen for response from all in the format of TYPE:[type]:NAME:[friendly name].
This seems to be problematic as I do not get reliable responses. At most the gateway gets 1-2 nodes responding at a time.

Any advice would be appreciated.

Felix

What kind of hardware are you running on exactly and with what library?
There are some related posts on the forum that talk about how to control remote node IDs, see: http://lowpowerlab.com/forum/index.php/topic,277.msg1417.html#msg1417

overdrive

#2
I am using atmega328p-pu/au, atmega644, atmega1284, atmega32u4  based custom boards all running at 16mhz with the standard bootloaders. I am using the RFM69 library in high power mode.
The devices are sprinkler/pump controller, perimeter lighting control, ground moisture sensors, rgb led controllers, home lighting controllers, smart timers, geyser timer, access control and security (reporting only) node, wendy house lighting, electricity control with motion and door sensor, Gate open sensor.

These nodes talk with a gateway node and it communicates with my .net serial monitor gateway on my windows server. the server in turn transmits/receives the communication via websockets to windows/android clients and a website driven by jquery.

What I want to do is to enumerate the active nodes of specific type when we access the control/monitoring section of the software I wrote. Example: Access webpage for rgb light controllers -> the dropdown list is populated with the active rgb controllers' friendly names and the values with the node IDs. Now you can select the one you want and start controlling it.

Everything is done except automatically enumerating the devices that are plugged in and ready to be commanded.

Thank you for the link. I use a similar method for the storing of setting the various settings of my devices in eeprom. I save the node_id, power level, friendly name, network, encryption key, type, available functions, version and device specific settings like timers and sensor calibration.

I was hoping there is an easy way to poll (one call) all my devices and get response of who is out there, powered on, their type and friendly name. I am aware this can be done with a database and a server routine to poll each possible id on the network and update the database with the response, but it will cause too much chatter and is overkill for what I want?

Felix

Quote from: overdrive on January 15, 2014, 04:14:55 PM
... with the standard bootloaders
There's the issue. Moteinos come with a custom bootloader that does the flashing: "Dualoptiboot"

overdrive

I dont see how. I am aware of the bootloader, but I am not interested in flashing the nodes remotely (not yet) Everything works perfectly except if I send out the "WHO" command. 10 devices try to communicate with the gateway node and only 1 or 2 gets through.

Felix

Ah, never mind, I got myself confused trying to respond to several threads at the same time.
I can only speak for atmega328p, not the others. You have to mind the fact that when you're asking the nodes to speak, they will all try to do so at the same time. The library has some leverage for that in that it listens for RF traffic and only allows a node to transmit once the channel is clear. But that's not bullet proof. I think you might still get collisions. I would try to add some random timing to the nodes before they respond to a WHO. Like wait random(100ms) and respond then, something like that...

jarrods

overdrive,
I can second what Felix is saying. If you send a command out that multiple nodes will respond to you will get mixed results unless they respond in a staggered fashion. for example:"delay(random(10,150);". You can also use "sendWithRetry" but from what I have seen that only works with a low number of nodes otherwise it chokes the network while everyone is trying to respond (albeit for less then a second), but the network is unusable during that time and you don't get a full report unless you set 5 or 10 retries. very messy.

In my case what I am doing is maintaining a list of nodes that have been configured in the eeprom of the gateway arduino. Then when I call for a status check via serial it will go through the list (stored in eeprom)and check them individually. I like this approach because that way you can still listen for events on your network while doing a network check. IE check node 1... listen for traffic ... check node 2.... listen for traffic... ect.

overdrive

Quote from: Felix on January 15, 2014, 06:44:19 PM
Ah, never mind, I got myself confused trying to respond to several threads at the same time.
Lol, figured as much when I saw the other threads.

Thanks guys.
I came to the same conclusion but wanted to ask just in case I missed anything. The sendwithretry does get very messy when more than 2 nodes try to reply. Collisions are inevitable it would seem  :-\ , now to plan for them and mitigate. I will give the random response time a shot, otherwise I will have to maintain a database with sporadic "are you there" checks.

LazyGlen

Could the Node ID be used to help out in this regard? Assuming for the moment that you ID'd the nodes starting at 1 and increment from there:

DEFINE QUE_DELAY 100 //average number of milliseconds an uncontested transmission and acknowledge requires + fudge factor
GATEWAY: send_to_all("Hey, tell me who's out there!")
NODE: delay(NODEID * QUE_DELAY); send_to_gateway("What do you want NOW?!");

This allows all nodes to use the same function for the roll call and keeps roll call as short as possible. If you use a random delay, you need to constrain it to a reasonable period, but as you add nodes, that period will increase, requiring either an update to the code or the gateway transmitting what it thinks a reasonable period is.

LG

KanyonKris

#9
This is an interesting topic, wireless architecture.

I like Glen's approach as it adds a fairly simple tweak to your current "who's out there?" approach.

Some other ideas:

Instead of the gateway shouting "who's out there?", the gateway could just listen for a while and keep a list of all the remote IDs it hears. This works for nodes that are periodically transmitting (ie. reporting temperature, ground moisture, etc.). For nodes that don't report but sit and wait to be commanded to do something, you could add code to have them transmit "I'm here" when they start up until they gateway acknowledges them and adds them to its list of heard remotes.

If you know your node IDs are between 10-30 the gateway could ping each ID. Something like: "Node 10, you out there?", "Node 10 here". Node 11, you there?", "Node 11 here".

Once you have a list of Node IDs you can ping them every so often to make sure they're alive.

My first Moteino project only needed the remote node to send info, not listen for commands so I have it sleep a lot to save battery power. It sleeps for 1 minute, wakes up, looks around and if it has nothing to report it goes back to sleep. If something has changed it transmits that info when it wakes up, listens for a while (in case the gateway has a new sketch to send it, wireless reprogramming), then goes back to sleep. I will also be adding code to have it transmit a heartbeat "I'm here and OK" every 5 minutes if it hasn't reported a change.

overdrive

Thanks for the suggestions. I think in the interest of simplicity and scalability I will go with adding a small DB/file to my server and have the server poll each possible node (2-255) every hour or so. This will keep an active sheet and once any application tries to command or monitor a device I will just check again if the device is out there and turned on. This way I can centrally manage it all. This will also save power on nodes as they do not need to worry about waking up and shouting I am here unnecessarily.  This approach will give me greater control over how noisy (RF wize) my system is if it need to take it for certification.

@LazyGlen: Your approach is also quite genius in it's simplicity. It might be a partial answer to my question. I will give it a shot as part of the implementation above.
The reason it is not the answer is that it will become a problem the larger the network is. To add to this issue, I have logically sliced the network for certain devices, i.e: RGB controllers use id 90-120 * 100ms or even less will cause a 9-12sec delay in response and the original idea of "on the fly" network topology mapping will be very slow from an end-user perspective.  But it will save on rf calls as the server can call once and then just listen. Once I get time to implement this I will post my results.

LazyGlen

#11
Quote from: overdrive on January 17, 2014, 02:49:14 AM
...
The reason it is not the answer is that it will become a problem the larger the network is. To add to this issue, I have logically sliced the network for certain devices, i.e: RGB controllers use id 90-120 * 100ms or even less will cause a 9-12sec delay in response and the original idea of "on the fly" network topology mapping will be very slow from an end-user perspective.  But it will save on rf calls as the server can call once and then just listen. Once I get time to implement this I will post my results.

I can write pseudo code ALL DAY man, it's turning it into something a compiler can read is the problem.

So you have assigned NODEID slices to different types of devices, makes sense. By making the gateway a little smarter when it takes a roll call, we can reduce the time it takes to do so. The gateway knows the network architecture, and it should know roughly how many nodes are in that section.

DEFINE QUE_DELAY 100 //average number of milliseconds an uncontested transmission and acknowledge requires + fudge factor
INT START
INT END

GATEWAY: send_to_all("Nodes START to END, Roll Call!")
NODE:
if (START <= NODEID <= END)
{   
   x = NODEID - START;
   delay(x * QUE_DELAY);
   send_to_gateway("Present!");
}

Thus the first node in the section the gateway is requesting roll call from transmits "Present!" immediately, and subsequent nodes follow as fast as you can get QUE_DELAY to reliably work. If the gateway is maintaining a list of active nodes, once it has received responses from the last known unit in the section and more than ~(QUE_DELAY * 3) milliseconds have past, it's probably not going to get any more responses.

I think that having the gateway keep track of assigned NODEID's makes a sort of dynamic NODEID assignment possible, when combined with other posted ideas of putting the NODEID in flash. At programing time you program the flash with a NODEID = NEWBIE. (An INT value that all nodes and gateways in your architecture recognize. I would use either the max valid number for NODEID, or perhaps better, any NODEID < the Gateways NODEID. Thus the Gateway code can recognize special cases.) When you power up a  node (one at a time if you only have 1 available NEWBIE id, or several if using the NEWBIE < GatewayID method, depending on where you set your GatewayID.):

NODE:
NODEID = Read.flash (location_node_id);
if NODEID == NEWBIE;
{
   send_to_gateway("Newbie online. NODE_TYPE");
   listen_for_response(NEW_NODE_ID);
   write.flash(location_node_id, NEW_NODE_ID);
   NODEID = NEW_NODE_ID;
}


When the Gateway hears a newbie come online, it can either look up what nodes are active for that NODE_TYPE, or do a roll call for the nodes set aside for that NODE_TYPE, then transmit back to the newbie the NODEID that it should be using. Assigning the first available NODEID allows nodes to be taken out of service and keep the list short, but the wetware may not know that a node went out of service (battery failure). Using next node in the list for that node type lets the wetware keep better track of the new nodes, but leaves empty NODEID's in the list when nodes go out of service, increasing the length of time for a roll call, and running the risk of aborting roll call if the code is not robust enough (eg: if several consecutive nodes go out of service.)

LG

thinkpeace

The idea of sending a who's online broadcast is interesting.  A quick way to enumerate all the modules that are operating within range.  This would be useful if the gateway needs to quickly determine all the nodes that are online.

Is there some reason that you would need to determine very quickly which nodes are online?

This is the approach that I use

Sleepy module online status update

On my RF network, my modules periodically send status messages to the gateway, and the gateway keeps track of which modules are online.  This way the modules can sleep between status updates and save on battery power.  The gateway has a timeout setting which is used to determine which modules are online or offline.

Automatic node number assignment

I came up with a way that modules are automatically assined node numbers.   I assign a unique long ID to each module when I program it, which is reported with the status messages.  The gateway keeps in SD Flash a list of what node number each module uses.  When a node is added to the network, it initially will have a random node number.  When it reports to the gateway, if the node number is already in use by another module, the gateway will reply assigning a new node number.  The sensor modules know not to sleep for a brief period after sending a sensor message.  The sensors also know not to reply to messages that are addressed to a different long ID.   

Messages from the gateway to the modules

Whenever the sensor modules send a status updates, before going back to sleep none or more of the following can happen.

  • node reassignment is received from the gateway
  • digital output command is received from the gateway
  • set status update frequency message is recevied from the gateway


KanyonKris

thinkpeace, cool that you have all that working. Please consider showing us your code. I for one would love to see how you did it and incorporate your techniques.

I like your way of handling new nodes. I had a similar idea:

The gateway is always ID 1. New nodes are ID 5 - 9. Established nodes are 10 - 250. When the gateway hears a new node, say ID 5, it replies back with the next highest unused ID in the 10 - 250 range (say IDs 10, 11, 12, 13 are in use, ID 14 would be issued) and the new remote node switches to the issued ID.

thinkpeace

KanyonKris, Your node assignment algorithm sounds similar to what I will be doing. 

I will upload it when it's finished.  As of yet I have the module status update working.  I still need to finish the node assignment and the IO output commands. 

I'm using the ChipKit WF32 as the Arudino Ethernet didn't have enough memory.  As written it would require something with much more than 2K of RAM.  It could be a ChipKit, AT2560Mega, Arduino Due, Raspberry Pi or Netduino.

Another option would be to implement the logic on a server running on a PC, or hosted by a web-host.  Then a Arduino with 2K of RAM would suffice.  The Arduino would simply relay messages back and forth to a server.  The server would implement the logic as I described.  I thought about this option, but I want to add more intelligence to the gateway, so it can make decisions and turn things on and off according to automation plans - even if the server is offline.  With the gateway making these decisions, it needs to keep track of the modules without depending on the server.  Thus I opted for a controller with more program memory and SRAM.

To interface with the server I am using JSON messages.  As I don't have a static IP address on my home network where the gateway resides, all communication with the server is initiated by the gateway.  The server will que up messages that need to go to the gateway, and send them in response to an update received from the gateway.